A ComfyUI workflow for transforming image styles or making targeted edits with Qwen Image 2.1, while guiding the model to preserve the rest of the scene.
Features
Style mode: choose from 14 editable presets—Realistic, Anime, Anime Flat Style, Semi Realistic Game 3D Style, Oil Painting, Illustration, LEGO, Medieval, 90/2000’s Cameras, Pixel Art, Pop Art, Japanese Manga, Korean Manhwa and Chinese Manhua—or write a prompt with Custom.
Edit mode: describe the requested change and combine it with an editable preservation prompt.
Four image inputs: Image 1 is the base image; optional reference slots 2–4 start bypassed.
Optional Face and Eyes refinement: separate passes with semantic face masks, a mask preview mode and a report explaining detected or skipped regions.
Manual LoRA loading: the rgthree Power LoRA Loader is not controlled by the style presets.
Getting started
Install the included qwen_reference_detail custom node and its requirements, then load the workflow JSON in ComfyUI. You’ll also need a working Qwen Image 2.1 setup and rgthree-comfy. On first use, the face detailer downloads Grounding DINO Tiny and a face-parsing model.
In Style mode, select a preset and adjust its prompt or add optional scene details. In Edit mode, enter the requested change; the workflow appends the preservation prompt before sending it to Qwen.
Face and Eyes refinement can be switched off independently. Use Mask only to preview eligible regions before applying the detail passes. Automatic masks can fail on small or highly stylized faces, so inspect the preview and final image.
Optional LoRAs
LoRAs are optional and remain manually controlled. Check each model page for compatibility, trigger words and recommended weights. The workflow works without them.
Credits
rgthree-comfy Power LoRA Loader
Face Parsing by jonathandinu
Review the face-parsing model’s license and usage terms before use. The mask system reduces unwanted edits but cannot guarantee pixel-perfect boundaries or identical results on every image.















