An outpaint LoRA for Qwen Image 2.1: pad a picture with flat gray on any side, and it fills the gray with more of the scene while the picture stays exactly where it was. One side, two, a corner, all four, a small picture in the middle of a big canvas, or a tilted picture with gray wedges in its corners.
Sometimes Qwen Image 2.1 outpaints just fine on its own, and sometimes it needs help. I tested 8 photos neither version was trained on, cropped so the original is the answer key, with 2 seeds and the same prompt with and without the LoRA.
Plain Qwen Image 2.1 moved or rescaled the picture in 13 of 16 renders (up to 284 px), and 5 of those showed a cut or ghosted seam after the original was pasted back.
With v1 or v2 the picture stayed in place in all 32 renders, and no seam showed.
When plain Qwen happens to keep the picture in place, the LoRA gives about the same result.
Which version
v2.0: everyday extensions at 1-2 MP. Trained on a people- and outfit-heavy set, and a little better than v1 on small and medium extensions.
v1.0: any extension size, including big zoom-outs, at about 1 MP.
Both fix the same thing and measure almost the same.
How to use it
Pad the picture with flat gray
#808080, each side a multiple of 32: about 1 MP for v1, 1-2 MP for v2. For a tilted picture, use Image Crop + Rotate + Pad with feather 0 (see below).Encode: give the padded canvas to Text Encode Qwen Image 2.1 as
image_1, at resolution 0, so the reference and the output share the canvas size.Prompt: the trigger instruction first, word for word, then optionally
Scene:and a description of the picture.Sample on the encoder's latent: 25 steps, CFG 1,
euler/simple, denoise 1, LoRA at strength 1.Decode and stitch: Qwen 2.1's VAE decodes RGBA, so run Split Image with Alpha, then paste the original pixels back over the result (Stitch Inpaint in my nodes).
Don't pin the known area with Set Latent Noise Mask: on Qwen 2.1 that draws a visible rectangle at the seam, and the LoRA keeps the picture in place on its own.
Easiest way: my free Qwen Image 2.1 Outpaint workflow does all of this. Qwen3-VL writes the description for you in the same style as the training captions, and there's a line for extra detail (for example, what she's wearing below the crop). It loads v2 by default and stops with a clear message if the LoRA file is missing, so you never run plain Qwen 2.1 by accident.
Tilted pictures
Straightening a crooked photo leaves gray wedges in the corners instead of straight borders. Neither version was trained on diagonal borders, and both handle them.
I rotated three pictures 17° and ran 36 renders (two seeds, with and without feather). Without the LoRA, Qwen 2.1 moved the scene in all 12 of its renders, by up to 95 px near the edges. The pasted-back original then doesn't line up: a lighter tilted box, cut or doubled edges, ghosted objects along the tilt. With v1 or v2, the picture stayed exactly in place in all 24.
My Qwen Image 2.1 Rotate + Outpaint workflow does the turn, crop and pad in one node with the right settings. Doing it by hand, set feather to 0 on Image Crop + Rotate + Pad. Its feather also fades the picture itself into the gray, which the model paints as a darker band along the tilt. The stitch still blends the edge.
What it fixes
On held-out pictures (never trained on), rendered in ComfyUI with the INT8 model, 25 steps, CFG 1:
Picture stays in place (kept-area PSNR):
subtle crops: 16.4 dB without the LoRA, 34.6 dB with v1, 34.7 dB with v2
big crops: 25.5 / 33.9 / 34.0 dB
Gray left unfilled (v1, 12 pictures): 9.1 % without, 0.9 % with.
Colour step at the seam (v1): +2.2 without, −0.6 with.
The full grids are on the Hugging Face model card. With a detailed prompt, plain Qwen 2.1 gets closer. Through my workflow, with its auto description, the new areas still landed closer to the real pictures with the LoRA (v1 fill error 23.3 vs 26.8).
Training
ai-toolkit (
qwen_image_2) on the Comfy-Org INT8 convrot base, the same weights ComfyUI runs. Rank 32, alpha 32, AdamW8bit, learning rate 1e-4, batch 1.v1: 1500 steps.
924 pairs from 231 pictures, each cropped four ways, drawn from six layouts: all sides, one side, opposite sides, a corner, three sides, and a small window keeping 12–30 %.
The target is the picture at up to 1 MP on the /32 grid. The source is the same canvas with everything outside the kept area painted
#808080.It works from step 500 and levels off around 1250.
v2: 2000 steps.
68 hand-picked pictures: 47 of people and outfits, 21 of scenes, interiors, anime and paintings.
Smaller extensions: the kept picture covers 45–93 % of the canvas.
Targets at 1-2 MP.
Captions: the trigger instruction (70 %) or a paraphrase (30 %). For 75 % of pairs,
Scene:plus a Qwen3-VL-8B description of the whole picture.Pictures: my own set plus openly licensed images:
Flickr photos from CommonCatalog CC-BY
museum art from PD12M (CC0 / public domain)
anime from anime-with-caption-cc0
Unsplash photos under the Unsplash License (v2)
Limits
Very large extensions (the kept picture under about 15 % of the canvas) leave a lot to invent, and results vary more by seed. Extend in two passes, or use v1.
A flat backdrop (a plain studio wall, a clear sky) can get a soft vertical smudge on the new side.
On shallow-focus photos, tilt wedges come out a little softer than the in-focus subject: the model continues the blur it sees.
v1 was trained at about 1 MP and v2 at 1-2 MP; larger canvases are untested.
Text in the new area is plausible, not legible.
About the gallery
The cover video widens a portrait to a square in one pass with v2: the gray grows out on both sides, then the new area is revealed.
The picture keeps the middle of a portrait and grows it back on all four sides with v2. It's a straight output of my Qwen Image 2.1 Outpaint workflow: drag the PNG into ComfyUI to load its exact settings.
Both portraits were made with Krea 2, and the LoRA never saw them in training. The with/without comparison grids are on the Hugging Face card.
License
A LoRA for Qwen Image 2.1, which is released under the Qwen Research License: non-commercial use only. Using the base model, and this LoRA with it, follows that license. Also on Hugging Face: ausboss/Qwen-Image-2.1-Outpaint-LoRA, with extra v1 checkpoints.
Changelog
v2.0 (2026-09-26): trained at 1-2 MP on a people- and outfit-heavy set (step 2000).
v1.0 (2026-09-25): first release (step 1500).
Questions or bugs: the comments here or GitHub issues. Post what you make with it; I read everything. I post new workflows on X @Zanzibased and GitHub.
Description
Everyday extensions at 1-2 MP, trained on a people- and outfit-heavy set (step 2000); a little better than v1 on small and medium extensions. Same file as qwen-image-2.1-outpaint-v2.safetensors on Hugging Face.
