Qwen Image 2.1 in ComfyUI: Local Setup by VRAM
A step-by-step ComfyUI guide to choosing official Qwen Image 2.1 weights by VRAM, saving a first PNG, diagnosing OOM and missing-model errors, and deciding when GGUF is worth trying.
Contents

The most reliable way to run Qwen Image 2.1 on a consumer GPU is not to begin with the smallest GGUF you can find. Update ComfyUI, load the official text-to-image workflow, select the official INT8 diffusion model plus a matching Qwen3-VL text encoder and VAE, then generate and save one small image. On an 8 GB or 12 GB card, do not start with BF16, 2K output, or the prompt enhancer.
Neither Qwen nor Comfy Org publishes one minimum-VRAM number that applies to every PC. The 8 GB and 12 GB configurations below are conservative starting points, not guarantees: system RAM, model offloading, drivers, resolution, and other GPU processes all change peak memory use.
Choose the weights first: file size is not peak VRAM
As of 2026-09-28, the official Comfy Org model repository contains two diffusion weights, three main text encoders, one VAE, and two separate prompt-enhancer models. A first image needs the diffusion model, one main text encoder, and the VAE. The qwen3.5_9b_..._pe_t2i file is an optional prompt enhancer; it does not replace the qwen3vl_8b_... encoder.
| Role | Recommended first-run file | Approx. download size | Folder |
|---|---|---|---|
| Diffusion model | qwen_image_2.1_int8_convrot.safetensors | 7.26 GB | ComfyUI/models/diffusion_models/ |
| Main text encoder, lower-memory choice | qwen3vl_8b_w4a8.safetensors | 6.31 GB | ComfyUI/models/text_encoders/ |
| Main text encoder, official template default | qwen3vl_8b_int8_convrot.safetensors | 9.35 GB | ComfyUI/models/text_encoders/ |
| VAE | qwen_image_2.1_vae_bf16.safetensors | 0.68 GB | ComfyUI/models/vae/ |
| Optional T2I prompt enhancer | qwen3.5_9b_qwen_image_2.1_pe_t2i.int8_convrot.safetensors | 9.47 GB | ComfyUI/models/text_encoders/ |
Those figures are repository file sizes, not runtime VRAM. INT8 diffusion + W4A8 encoder + VAE is about 14.24 GB of downloads; the official template’s INT8 + INT8 + VAE combination is about 17.28 GB. Enabling the prompt enhancer adds about 9.47 GB of storage and another large model-loading stage.
A practical first configuration for each VRAM tier
| Available VRAM | Try first | First-generation settings | Upgrade only after success |
|---|---|---|---|
| 8 GB | INT8 diffusion + W4A8 main encoder + VAE; prompt enhancement off | 768×768 or below roughly 1 MP, batch 1, CFG 1 | Try 1024 next; consider community GGUF only if the official path still OOMs |
| 12 GB | INT8 diffusion + W4A8 main encoder; try INT8 encoder later | 1024×1024, batch 1, CFG 1 | Enable the enhancer or raise resolution one change at a time |
| 16 GB and up | Official-template INT8 diffusion + INT8 encoder + VAE | 1024×1024, 25 steps, CFG 1 | Test prompt enhancement and 2K separately |
| Plenty of VRAM and system RAM | Prove the INT8 path first, then compare BF16 components | Keep prompt, seed, and dimensions fixed | Use BF16 only when you accept the larger footprint and want a controlled comparison |
The 8 GB and 12 GB rows minimize troubleshooting cost; they are not official hardware requirements. Qwen documents model offloading for limited-memory GPUs, while ComfyUI decides what to keep on the GPU for a particular machine. Two systems with the same GPU can therefore behave differently.
1. Update ComfyUI before looking for custom nodes
Qwen Image 2.1 has native ComfyUI support. The official workflow should not require a third-party node pack. Red nodes, a missing Qwen Image 2.1 template, or an absent TextEncodeQwenImage21 node usually mean the core application or its packaged dependencies are out of date.
- Windows Portable: close ComfyUI and run
ComfyUI_windows_portable/update/update_comfyui.bat, then restart. Do not useupdate_comfyui_and_python_dependencies.batas the routine first choice because it reinstalls the dependency stack. - ComfyUI Desktop: update the Engine from Manage → Update. Stable releases can lag the newest model support; use the available
Latest on GitHubchannel, or move to Portable/manual installation if your current Desktop build still lacks the nodes. - Manual Git install: update code and dependencies inside the Python environment that runs ComfyUI, then restart it.
cd /path/to/ComfyUI
git pull
python -m pip install -r requirements.txt
python main.py
The official update guide warns that git pull alone can leave the frontend, workflow templates, and new nodes outdated. Updating requirements.txt dependencies is part of the update, not an optional cleanup step.
2. Put each model in the folder its loader scans
Download your selected files from the official ComfyUI model repository. Keep every filename unchanged; browser-added suffixes such as (1) can prevent the workflow from finding the model it names.
A low-memory first run can use this layout. Choose either the W4A8 or INT8 main encoder:
ComfyUI/
└── models/
├── diffusion_models/
│ └── qwen_image_2.1_int8_convrot.safetensors
├── text_encoders/
│ ├── qwen3vl_8b_w4a8.safetensors
│ └── qwen3.5_9b_qwen_image_2.1_pe_t2i.int8_convrot.safetensors # optional
└── vae/
└── qwen_image_2.1_vae_bf16.safetensors
To match the official template’s default, replace qwen3vl_8b_w4a8.safetensors with qwen3vl_8b_int8_convrot.safetensors. Restart ComfyUI after copying the files; loader dropdowns normally reflect the files found during startup.
3. Import the official text-to-image workflow
Open the Qwen Image 2.1 text-to-image template from ComfyUI’s Templates panel, or download and drag the official T2I workflow JSON onto the canvas. The current official template packages the main graph as a subgraph and starts at 1024×1024, 25 steps, CFG 1, and euler + simple.
Check the loaders and controls in this order:
- Select
qwen_image_2.1_int8_convrot.safetensorsas the diffusion model. - Select
qwen3vl_8b_w4a8.safetensorsorqwen3vl_8b_int8_convrot.safetensorsas the main text encoder, with typeqwen_image. - Select
qwen_image_2.1_vae_bf16.safetensorsas the VAE. - Keep
refine_promptset tofalsefor the first run. The current template exposes a PE-model selector, but theqwen3.5_9b_..._pe_t2imodel is only for prompt enhancement. - Start an 8 GB card at 768×768. Start 12 GB and larger cards at 1024×1024. Prefer dimensions divisible by 32.
- Keep batch size 1 and CFG 1. Do not add LoRA, ControlNet, reference-image, or upscaling branches yet.
Use a deliberately simple first prompt so that success is easy to judge:
A red ceramic teapot on a wooden table, soft window light, plain background.
Do not begin at 2K. Native 2K support means the model can generate 2048×2048 directly; it does not mean an 8 GB or 12 GB GPU can complete that size under every configuration.
4. Queue one job, wait, and verify that the PNG was saved
Click Queue Prompt once. The first run may load tens of gigabytes and move components among VRAM, system RAM, and the CPU. If the canvas has no new preview yet, watch the launch terminal or log rather than repeatedly adding the same job to the queue.
Treat the first run as successful only when all four signals are present:
- The queue finishes without
model not found,missing node type,CUDA out of memory, or a crashed process. SaveImageAdvanceddisplays the image; the sampler finishing by itself is not enough.- A PNG exists under
ComfyUI/output/. The official template uses the filename prefixQwen_image_2.1by default; if you changed the Save Image node, use the location shown there. - The PNG opens normally, has the expected dimensions, and is not fully black, fully transparent, or corrupt.
On an NVIDIA system, you can watch memory in a second terminal:
nvidia-smi -l 1
Record the diffusion file, text encoder, dimensions, and observed peak. When you later change weights or resolution, change one variable at a time so a new failure has an identifiable cause.
5. Troubleshoot by symptom, not by replacing everything
A node is red or says missing node type
Update ComfyUI core and dependencies, then restart. TextEncodeQwenImage21, ResolutionSelector, and the logic switch in the official workflow are core ComfyUI nodes. Installing an unrelated third-party pack with a similar name can make the graph harder to diagnose.
A model dropdown is empty or reports model not found
Check the folder, extension, and exact filename, then restart ComfyUI. Common mistakes are placing the main encoder in models/clip/, placing the diffusion model in models/checkpoints/, or allowing the browser to rename a duplicate download.
VRAM runs out before sampling starts
Turn prompt enhancement off, switch the main encoder from INT8 to W4A8, keep batch 1, and close GPU-heavy browsers, games, or video tools. Drop an 8 GB card to 768×768; return a 12 GB card to 1024×1024 and remove optional branches.
VRAM runs out during VAE decode or just before saving
Reduce width and height, then queue one new job. Lowering steps mostly changes sampling time; reducing resolution more directly lowers latent and decode memory.
The queue looks frozen while disk or CPU activity continues
This is often model loading or offloading rather than a finished failure. Wait for a clear completion or error state. Repeated clicks only stack more large jobs. If system memory keeps climbing into heavy swapping, cancel, lower the configuration, and restart ComfyUI.
The image looks washed out, overcooked, or structurally broken
Confirm that CFG is still 1. The official template notes that the negative prompt is unused at CFG 1; copying a high-CFG habit from an SDXL workflow is a poor starting point here.
1024 works but 2K always OOMs
Keep 1024 as the working baseline. Try 1280, then 1536, and only then 2048, changing one tier at a time with batch 1 and prompt enhancement off. Model capability and local memory capacity are separate constraints.
What 8 GB and 12 GB users should actually choose
On 8 GB, the goal is to prove the pipeline, not to enable every template feature. Use the INT8 diffusion model, W4A8 main encoder, and VAE; disable PE; start at 768×768 and batch 1. If this still fails, confirm that ComfyUI is current and that the machine has enough system RAM before moving to GGUF. Do not insert an unknown quantized file into the stock workflow and assume it is interchangeable.
On 12 GB, use 1024×1024 as the first target. Start with the W4A8 encoder for headroom, then try the official template’s INT8 encoder after the base run is stable. Prompt enhancement is a second-stage feature: it loads another 9B model to rewrite the prompt before image generation, so it should not be introduced while the base graph is still failing.
Community posts only show that individuals are experimenting; they do not establish a minimum. One user said they were running the model on an 8 GB RTX 5060 laptop while trying to optimize it. A separate 12 GB RTX 3060 user suggested W4A8 to an 8 GB user. These are the first report and the second suggestion, not controlled tests with complete logs, so they are useful as leads rather than guarantees.
GGUF is an optional low-memory route, not the official first run
The Qwen and Comfy Org templates use .safetensors. GGUF requires a third-party loader, a Qwen Image 2.1-specific quantized file, and a graph built for that loader. The ComfyUI-GGUF project describes itself as work in progress and does not certify every third-party Qwen quant.
Consider the community route only after the official INT8/W4A8 path still fails for memory:
- Update ComfyUI, then install
ComfyUI-GGUFthrough Manager or from its repository. - Obtain a file explicitly built for Qwen Image 2.1. Check the publisher, license, hash, and required loader.
- Replace the stock diffusion loader with
Unet Loader (GGUF). Keep a Qwen Image 2.1-compatible main text encoder and VAE. - Select the exact filename in the loader and test below roughly 1 MP, CFG 1, batch 1.
- When it fails, return to that quant’s own workflow notes. Do not mix a Qwen Image 1.x, Flux, or generic Stable Diffusion graph into the diagnosis.
A Russian user reported getting qwen-image-2.1-UC-Q4_K_M.gguf to run in ComfyUI but also said the installation required extra help. A separate community Q8_0 workflow combines GGUF with the official text encoder and VAE. Those reports show that a community route exists; they do not prove that arbitrary Q4 or Q8 files are compatible or safe downloads.
Add prompt enhancement, 2K, and editing only after the baseline works
Once the base image is stable, extend the graph in this order:
- Prompt enhancement: download
qwen3.5_9b_qwen_image_2.1_pe_t2i.int8_convrot.safetensors, select it, and turn onrefine_prompt. Keep seed, dimensions, and steps fixed when comparing the result. - Higher resolution: raise dimensions in stages instead of jumping from 1024 to 2048×2048. Save as PNG when you need an alpha channel.
- Image editing: switch to the official Image Edit workflow. Its prompt enhancer uses the
..._pe_i2i...file, not the T2I file. - Transparent output: use the official RGBA/transparent-background prompt wording and save as PNG. First prove an ordinary opaque image so transparency is the only new variable.
Final acceptance checklist
- ComfyUI core, frontend dependencies, and workflow templates are current; no core node is red after restart.
- The diffusion model, main text encoder, and VAE are in their correct folders and visible by exact filename.
- Prompt enhancement is off, batch is 1, and CFG is 1; 8 GB starts at 768, while 12 GB starts at 1024.
- Only one task is queued, and the terminal finishes without a missing-model, missing-node, or OOM error.
- A valid PNG exists under
ComfyUI/output/with the requested dimensions. - The successful weight combination, resolution, and observed peak memory are recorded before PE, higher resolution, editing, or GGUF is added.
When all six are true, you have a reproducible local Qwen Image 2.1 baseline. Every later optimization should change one element of that successful baseline, not replace it with a large graph whose failure source is impossible to isolate.