Mage Flow is a 4-billion parameter image generation and editing stack from Microsoft. It pairs a lightweight latent tokenizer (Mage-VAE) with a native-resolution multimodal diffusion transformer (NR-MMDiT), and it generates natively from 512 to 2048 pixels at any aspect ratio - no bucket quantization, no padding. At 4B it scores 0.90 on GenEval, the best result among open-source models, against systems five to eight times its size.
Originally released by Microsoft on Hugging Face under the MIT license. All credit for the model goes to the Mage-Flow team. Civitai is hosting a mirror so creators can run it on-site - please head to the original repository for weights, code, and updates, and cite the paper if you build on it.
Built by
- Microsoft - Mage project team, lead author Xinjie Zhang and contributors
- Paper: Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing (arXiv:2607.19064)
Native-resolution generation
NR-MMDiT processes variable-length image and text sequences with per-sample 2D rotary embeddings and joint self-attention. Because resolution is native rather than bucketed, extreme aspect ratios like 4:1 (512x2048, 2048x512) come out clean instead of stretched or cropped, and you are not locked to a fixed training resolution.
Instruction-based editing
The Edit variants handle semantic content changes, appearance transformation, image restoration, and structure-aware edits in one architecture - you describe the change in plain language rather than masking. On ImgEdit-Bench the RL-aligned edit model scores 4.34, with GEdit-EN at 8.127 and GEdit-CN at 8.123.
Versions on Civitai
All six upstream checkpoints are mirrored here:
- Mage-Flow-4B-Base - text-to-image, 30 steps, CFG 5.0
- Mage-Flow-4B - text-to-image, RL-aligned, 20 steps, CFG 5.0
- Mage-Flow-4B-Turbo - text-to-image, 4-step distilled, CFG 1.0
- Mage-Flow-Edit-4B-Base - image editing, 30 steps, CFG 5.0
- Mage-Flow-Edit-4B - image editing, RL-aligned, 30 steps, CFG 5.0
- Mage-Flow-Edit-4B-Turbo - image editing, 4-step distilled, CFG 1.0
If you are starting out, use Mage-Flow-4B for text-to-image and Mage-Flow-Edit-4B for editing. The RL-aligned builds are the quality sweet spot. Drop to the Turbo builds when you want speed and can accept a small quality tradeoff.
Runs light
Mage-VAE uses roughly 12x fewer encode and 22x fewer decode MACs per pixel than FLUX.2-VAE at matching reconstruction fidelity. Combined with native-resolution packing, which runs the conditional and unconditional CFG branches in a single packed forward pass, Mage-Flow-Turbo produces a 1024x1024 image in 0.59 seconds on a single A100 and peaks at 18-20 GB - the lowest memory footprint among the systems it was compared against.
Links
- Source: huggingface.co/microsoft/Mage-Flow-Base
- Code: github.com/microsoft/Mage
- Project page: microsoft.github.io/Mage
- Paper: arXiv:2607.19064
- License: MIT
Description
FAQ
Comments (10)
So, a new Open source model from Microsoft emurges. I wonder how this edit model would compare to Flux 2 Klein 9B or editing workflows of Krea 2. Would like to know.
It will make the bad end as Lens from Microsoft too...
The quality of the supplied base model is rather questionable, as it seems to have been overtrained on powerpoint text splashes and website screenshots.
However, the actual model architecture is super fast, the VAE is tasty, and the fact that it comes with an edit module out of the gate is nice.
It also shares a text encoder with Krea2, so training it on Krea2 output for a given prompt seems like a very good starting point - as that model produces very nice outputs, though it is comparatively glacial in its sampling process.
Like MS Paint to Photoshop.
CivitAI refuses to edit NSFW content with this.
The reason this model exists is the same as why winblows does. Waste of traffic and GPU time.
It is uncensored! And image quality a bit better than zimage IMHO. But spelling is worse.
Only boobs / blood are not censored. All other things don't work.
It feels like the model lay too much work on VAE. Results have very bad details. Yes, the model is fast as for its size but it not worth it. It looks even worse than a bad 2x upscale.
It is good if you restore skin details via Composite nods (you need it for almost any edit models specially if you want the face to remain unchanged), specially good for removing objects from images. its really fast, like ~10second on a 8GB card for editing an image near 2k resolution.
Krea 2 even tho it talke more than 1 min to edit and around 2 min for mixing two images and its not even an edit model is the only good model (+identity edit lora) for editing and awesome for generation.
Havent tried Gwen edit but i saw the result, close to mage flow with plastic looking skin.
Flux is just garbage both for generation and editing. that model was the king of all these plastic skins and double chin faces. SDXL inpaint is better at skins and faces than Flux.
Z Image Turbo is awsome, cant wait for its edit model. next to Krea 2 has the best skin textures and the best images generation and its fast.
Details
Files
mage_flow_vae_bf16.safetensors
Mirrors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
mage_vae_bf16.safetensors
diffusion_pytorch_model.safetensors
mage_vae_bf16.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
mage_flow_vae_bf16.safetensors
Mage-Flow-VAE.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
Mage-Flow-VAE.safetensors
Mage-Flow-VAE.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
mageFlow_mageFlowEdit4BTurbo.safetensors
Mirrors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
mage_flow_edit_turbo_bf16.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
diffusion_pytorch_model.safetensors
