Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -160,6 +160,7 @@ work_dirs
examples/regression_test/regression_outputs/
INSTALL_HYS.md
*.webp
!examples/data/qwen_image_21/*.webp
AGENTS.md
_version.py.mcp.json
telefuser/_version.py
Expand Down
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
116 changes: 113 additions & 3 deletions examples/qwen_image/README.md
Original file line number Diff line number Diff line change
@@ -1,13 +1,14 @@
# Qwen-Image Examples

These examples provide text-to-image generation, image editing, quantized inference, and feature-cache calibration
with Qwen-Image checkpoints.
with Qwen-Image checkpoints, including the unified Qwen-Image 2.1 pipeline.

## Model Source

| Model | HuggingFace | ModelScope | Purpose |
| --- | --- | --- | --- |
| Qwen-Image | [Qwen/Qwen-Image](https://huggingface.co/Qwen/Qwen-Image) | [Qwen/Qwen-Image](https://modelscope.cn/models/Qwen/Qwen-Image) | Text-to-image base weights |
| Qwen-Image 2.1 | [Qwen/Qwen-Image-2.1](https://huggingface.co/Qwen/Qwen-Image-2.1) | [Qwen/Qwen-Image-2.1](https://modelscope.cn/models/Qwen/Qwen-Image-2.1) | Native generation, editing, and reference examples |
| Qwen-Image-Lightning | [Qwen/Qwen-Image-Lightning](https://huggingface.co/Qwen/Qwen-Image-Lightning) | [Qwen/Qwen-Image-Lightning](https://modelscope.cn/models/Qwen/Qwen-Image-Lightning) | Distilled LoRA and FP8 variants |
| Qwen-Image-Edit | [Qwen/Qwen-Image-Edit](https://huggingface.co/Qwen/Qwen-Image-Edit) | [Qwen/Qwen-Image-Edit](https://modelscope.cn/models/Qwen/Qwen-Image-Edit) | Image editing |

Expand All @@ -17,21 +18,30 @@ with Qwen-Image checkpoints.
| --- | --- | --- |
| Text-to-image | Supported | BF16, Lightning LoRA, TeleFuser FP8, and NF4 examples |
| Image editing | Supported | TeleFuser and Diffusers reference paths |
| Qwen-Image 2.1 reference generation | Supported | One to ten condition images in the CLI example |
| Multi-GPU inference | Supported | CFG and Ulysses parallelism on scripts exposing `--gpu_num` |
| LoRA | Supported | Lightning LoRA example |
| Quantization | Supported | Pre-quantized FP8, online TeleFuser FP8, and NF4 |
| CPU offload | Supported | Used by native Qwen pipelines |
| Feature cache | Supported | Separate T2I and edit calibration tools |
| Server API | Supported | Native examples expose standard pipeline functions |

Qwen-Image 2.1 uses a 64-channel latent VAE, Qwen3-VL prompt encoding, and a
single-stream block-causal transformer. The native TeleFuser path uses 40 steps
by default with classifier-free guidance disabled. The same stage pipeline handles
text-to-image, image editing, and reference-guided generation.

## Requirements

- GPU: one H100-class CUDA GPU for the documented configurations; use supported multi-GPU degrees as needed
- Software: the standard TeleFuser installation; Diffusers is required for the official reference scripts
- Input assets: a readable image for editing; T2I requires no input asset
- Software: the standard TeleFuser installation plus Diffusers standalone Qwen-Image 2.1 VAE classes and Transformers Qwen3-VL classes
- Input assets: T2I requires no input asset; edit uses `examples/data/edit2511input.png`, and reference generation
uses five official demo images under `examples/data/qwen_image_21/`

Install TeleFuser by following the [development setup](../../CONTRIBUTING.md#development-setup).

The TeleFuser entry point implements the 2.1 DiT locally. It uses Diffusers only for the standalone VAE class and Transformers for the Qwen3-VL text encoder and processor.

## Model Directory

```text
Expand All @@ -41,6 +51,12 @@ ${TF_MODEL_ZOO_PATH}/
| |-- vae/
| |-- text_encoder/
| \-- tokenizer/
|-- Qwen-Image-2.1/
| |-- transformer/
| |-- vae/
| |-- text_encoder/
| |-- processor/
| \-- scheduler/
|-- Qwen-Image-2512-Lightning/
| \-- Qwen-Image-2512-Lightning-8steps-V1.0-fp32.safetensors
|-- Qwen-Image-Edit-2509/
Expand All @@ -66,10 +82,45 @@ python examples/qwen_image/qwen_image_t2i_h100.py \

The command writes the generated image to `work_dirs/qwen-image.png`.

The 2.1 standard example uses the same regression prompt, negative prompt, seed, and default `16:9` aspect ratio
as the existing Qwen-Image example. These defaults are declared in the Python scripts. Both examples map `16:9`
to **1664×928** pixels.

Select physical GPU 3 by masking it into the process (it is then visible as
`cuda:0`):

```bash
CUDA_VISIBLE_DEVICES=3 python examples/qwen_image/qwen_image_21_t2i_h100.py \
--model_root /hhb-data/aigc/model_zoo/Qwen-Image-2.1 \
--output_path work_dirs/qwen-image-2.1.png
```

## Examples

### Text-To-Image

#### `qwen_image_21_t2i_h100.py`

This standard example loads Qwen-Image 2.1 through TeleFuser for text-to-image
generation.

```bash
CUDA_VISIBLE_DEVICES=3 python examples/qwen_image/qwen_image_21_t2i_h100.py \
--model_root /hhb-data/aigc/model_zoo/Qwen-Image-2.1 \
--prompt "A neon shop sign that reads QWEN IMAGE 2.1 in the rain" \
--output_path work_dirs/qwen-image-2.1.png
```

Key options:

| Option | Default | Description |
| --- | --- | --- |
| `--model_root` | `/hhb-data/aigc/model_zoo/Qwen-Image-2.1` | Diffusers model directory |
| `--gpu_num` | `1` | Qwen-Image 2.1 currently uses one GPU |
| `--aspect_ratio` | `16:9` | Uses the existing Qwen-Image size mapping; default is 1664×928 |
| `--height`, `--width` | unset | Optional overrides for the mapped output dimensions |
| `--output_path` | generated example name | PNG output path |

#### `qwen_image_t2i_h100.py`

```bash
Expand Down Expand Up @@ -126,6 +177,33 @@ python examples/qwen_image/qwen_image_edit_plus_h100.py \
--output work_dirs/qwen-image-edit.png
```

#### `qwen_image_21_edit_h100.py`

The native 2.1 edit example uses the existing Qwen-Image edit test image. Without
`--height` or `--width`, output dimensions follow the input image aspect ratio
at approximately one megapixel.

```bash
CUDA_VISIBLE_DEVICES=3 python examples/qwen_image/qwen_image_21_edit_h100.py \
--image_path examples/data/edit2511input.png \
--output_path work_dirs/qwen-image-2.1-edit.png
```

### Reference-Guided Generation

#### `qwen_image_21_reference_h100.py`

The default test uses the official Qwen-Image 2.1 ["Outfit styling (5 images)" demo
case](https://huggingface.co/spaces/Qwen/Qwen-Image-2.1/blob/main/examples/cases.json). Its prompt and five input
images are included in the script and `examples/data/qwen_image_21/` in the official order: model, down jacket,
Mary Jane shoes, handbag, and fur hat. Run it without extra arguments, or pass `--image_path` once per custom
reference image. The output aspect ratio follows the last reference unless dimensions are supplied.

```bash
CUDA_VISIBLE_DEVICES=3 python examples/qwen_image/qwen_image_21_reference_h100.py \
--output_path work_dirs/qwen-image-2.1-reference.png
```

### Cache Calibration

#### `qwen_image_cache_calibrate.py`
Expand Down Expand Up @@ -166,7 +244,39 @@ python examples/qwen_image/qwen_image_edit_plus_official.py \
This reference keeps the official fixed sampling parameters while making model, input, prompt, and output paths
explicit.

## Serving

The standard entry points expose `t2i` and `i2i` service contracts. The edit
example uses the existing `i2i` service task because `telefuser serve --task`
does not offer a separate `edit` value.
The existing service schema supplies one image through `first_image_path` for
`edit` and `i2i`; use the reference CLI for multiple images.

```bash
CUDA_VISIBLE_DEVICES=3 telefuser serve examples/qwen_image/qwen_image_21_t2i_h100.py \
--task t2i --port 8091
```

To serve either image-conditioned example, select its script with the `i2i`
task. For example:

```bash
CUDA_VISIBLE_DEVICES=3 telefuser serve examples/qwen_image/qwen_image_21_edit_h100.py \
--task i2i --port 8092
```

Upload the source image through `/v1/tasks/form` with `first_image_file`.

The service loads one pipeline replica on the selected GPU and writes each
request result to the service output directory.

## Configuration

Supported aspect ratios include `1:1`, `16:9`, `9:16`, `4:3`, `3:4`, `3:2`, and `2:3`. The exact
resolution mapping is defined in each entry point.

## Troubleshooting

If loading fails with `safetensors ... invalid JSON in header`, one or more
checkpoint shards are incomplete. Verify the four files under
`text_encoder/` and copy them again before starting the service.
125 changes: 125 additions & 0 deletions examples/qwen_image/qwen_image_21_edit_h100.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,125 @@
"""Edit an existing image with the native Qwen-Image 2.1 pipeline.

Usage:
CUDA_VISIBLE_DEVICES=3 python examples/qwen_image/qwen_image_21_edit_h100.py \
--image_path examples/data/edit2511input.png \
--output_path work_dirs/qwen-image-2.1-edit.png
"""

from __future__ import annotations

import os
from pathlib import Path

import click
import torch
from PIL import Image

from telefuser.pipelines.qwen_image import QwenImage21Pipeline
from telefuser.service.core.contract_templates import build_pipeline_manifest, build_task_contract_template

PPL_CONFIG = {
"name": "qwen_image_2.1_edit",
"model_root": os.path.join(os.environ.get("TF_MODEL_ZOO_PATH", "/hhb-data/aigc/model_zoo"), "Qwen-Image-2.1"),
"prompt": '这个女生看着面前的电视屏幕,屏幕上面写着"阿里巴巴"',
"image_path": "examples/data/edit2511input.png",
"seed": 42,
"num_inference_steps": 40,
}

PIPELINE_CONTRACT = build_pipeline_manifest(
pipeline_name=PPL_CONFIG["name"],
supported_tasks=["i2i"],
task_contracts={
"i2i": build_task_contract_template(
"i2i",
required_inputs=["first_image_path"],
parameter_overrides={
"prompt": {"default": PPL_CONFIG["prompt"]},
"seed": {"default": PPL_CONFIG["seed"]},
},
excluded_parameters=["negative_prompt", "resolution", "aspect_ratio"],
)
},
)


def get_pipeline(
parallelism: int = 1,
model_root: str = PPL_CONFIG["model_root"],
device: str | None = None,
) -> QwenImage21Pipeline:
"""Load the shared native Qwen-Image 2.1 pipeline."""

if parallelism != 1:
raise ValueError("Qwen-Image 2.1 example currently supports one GPU")
return QwenImage21Pipeline.from_pretrained(model_root, device=device or "cuda", torch_dtype=torch.bfloat16)


def run(
pipeline: QwenImage21Pipeline,
prompt: str,
image: Image.Image,
seed: int = PPL_CONFIG["seed"],
height: int | None = None,
width: int | None = None,
) -> list[Image.Image]:
"""Apply an editing instruction to one source image."""

return pipeline(
prompt=prompt,
image=image,
seed=seed,
height=height,
width=width,
num_inference_steps=PPL_CONFIG["num_inference_steps"],
)


def run_with_file(
pipeline: QwenImage21Pipeline,
prompt: str,
first_image_path: str,
output_path: str,
seed: int = PPL_CONFIG["seed"],
height: int | None = None,
width: int | None = None,
) -> dict[str, str]:
"""Save an edited image for the standard service entrypoint."""

with Image.open(first_image_path) as source:
images = run(pipeline, prompt, source.copy(), seed=seed, height=height, width=width)
destination = Path(output_path)
destination.parent.mkdir(parents=True, exist_ok=True)
images[0].save(destination)
return {"output_path": str(destination)}


@click.command()
@click.option("--gpu_num", default=1, type=int, show_default=True)
@click.option("--model_root", default=PPL_CONFIG["model_root"], show_default=True)
@click.option("--image_path", default=PPL_CONFIG["image_path"], type=click.Path(exists=True), show_default=True)
@click.option("--prompt", default=PPL_CONFIG["prompt"], show_default=True)
@click.option("--seed", default=PPL_CONFIG["seed"], type=int, show_default=True)
@click.option("--height", type=int, default=None)
@click.option("--width", type=int, default=None)
@click.option("--output_path", type=click.Path(path_type=Path), default=Path("work_dirs/qwen-image-2.1-edit.png"))
def main(
gpu_num: int,
model_root: str,
image_path: str,
prompt: str,
seed: int,
height: int | None,
width: int | None,
output_path: Path,
) -> None:
"""Run Qwen-Image 2.1 image editing."""

pipeline = get_pipeline(gpu_num, model_root)
result = run_with_file(pipeline, prompt, image_path, str(output_path), seed=seed, height=height, width=width)
print(f"Image saved to: {result['output_path']}")


if __name__ == "__main__":
main()
Loading
Loading