The term "new benchmark" in the AI image generation space has been overused, but since the arrival of Flux.1, even users of Midjourney and SDXL have started to take it seriously. In August 2024, Black Forest Labs—a company founded by core team members from Stable Diffusion—released the Flux.1 series of models. With stunning image quality and precise understanding of prompts, it quickly sent shockwaves through the AI art community.
What is Flux.1?
Flux.1 is a text-to-image model series developed by Black Forest Labs, trained on a scale of 12 billion parameters. It shows significant improvements over previous models in image quality, human anatomy, and detail rendering. More importantly, its ability to understand prompts far exceeds that of SDXL—if you say "left hand holding an apple, right hand holding a phone," it will actually draw exactly that, rather than giving you two hands holding apples or randomly assigning objects.
The core members of Black Forest Labs come from Stability AI (the parent company of Stable Diffusion), including key researchers like Robin Rombach. You can think of it this way: the same team that created SD has built a stronger model.
Three Versions for Different Needs
Flux.1 is not a single model but is divided into three versions based on use cases:
Flux.1 Pro The flagship commercial version, available only via API and cannot be deployed locally. It offers the highest image quality, particularly excelling in character details, lighting handling, and complex compositions. It is billed per request and is suitable for commercial scenarios with the highest quality requirements. The API can be accessed through platforms like Replicate and fal.ai.
Flux.1 Dev An open-source research version with parameters similar to Pro. Its quality is slightly lower but the gap is negligible. It can be deployed locally, run on platforms like ComfyUI and AUTOMATIC1111, and used commercially (subject to the Flux Non-Commercial License). This is currently the most widely used version in the community and serves as the foundation for numerous fine-tuned models based on Flux.
Flux.1 Schnell A speed-prioritized version; "Schnell" means "fast" in German. It generates an image in just a few steps on consumer-grade GPUs, making it several times faster than Dev. While the quality is relatively lower, it still far surpasses many other models. This version is fully open-source under the Apache 2.0 license and can be used freely for commercial purposes.
Why It Shook the Community
Prompt Adherence
This is Flux.1's most frequently cited advantage. Previously, when using SDXL, you might write a detailed description but only get partial results. Flux.1 has made qualitative improvements here; its fidelity to complex scene descriptions and specific detail requirements is significantly higher. Especially regarding text rendering—having AI accurately write specified text within an image has long been a difficult problem. Flux.1 has made breakthrough progress in this area, capable of rendering English letters with considerable accuracy (Chinese support is still improving).
Human Anatomy
"AI drawing hands" used to be a meme because early models often made mistakes with human hands and finger counts. Flux.1 has significantly improved the accuracy of human anatomical structures. While not perfect, it is a clear step forward compared to previous models.
Realism and Detail
The rendering quality of details such as skin texture, fabrictexture (texture), and lighting changes has taken the "realism" of images generated by Flux.1 to another level. Even Schnell, the speed-prioritized version, produces image quality that surprises many people.
How to Use Flux.1
Online Experience, No Deployment Required
- Replicate.com: Use Flux.1 Pro and Dev directly in your browser. It is billed per request, and you don't need to prepare your own GPU.
- fal.ai: Also provides online interfaces with fast speeds, supporting the Flux series.
- Hugging Face Spaces: There are many free demos of Flux.1 Schnell deployed here. They are slower but cost nothing.
Local Deployment via ComfyUI ComfyUI is currently the most mainstream local solution for using Flux.1. Requirements include:
- An NVIDIA GPU with at least 12GB of VRAM (for the full Flux.1 Dev version)
- Or an 8GB VRAM GPU to run the 4-bit quantized version
- Download model weights from Hugging Face and install the corresponding workflow in ComfyUI
The ComfyUI community has a large number of ready-made Flux.1 workflows that can be imported directly, so there is no need to build from scratch.
AUTOMATIC1111 / Forge The Forge branch of Stable Diffusion WebUI (A1111) offers good support for Flux.1. It has a more user-friendly interface and is suitable for users who do not want to tinker with ComfyUI nodes.
API Integration Developers can call Flux.1 via Replicate API or fal.ai API to integrate it into their own applications without managing GPU infrastructure.
Comparison with Other Models
vs SDXL: Flux.1 wins on almost all metrics—prompt adherence, detail quality, and human anatomy. SDXL's advantage lies in its massive community accumulation of LoRAs and fine-tuned models; Flux is rapidly catching up in this regard.
vs Midjourney: Midjourney still has its own style in terms of artistic feel and composition, especially with the art styles of MJ v6 that many users like. Flux.1 Pro's realism and detail rendering can compete with MJ. Moreover, Flux can run locally without relying on Discord, offering more control.
vs DALL-E 3: DALL-E 3 is integrated into ChatGPT, making it easy to use with good prompt understanding. However, its image style tends to be "clean" and sometimes overly cartoonish; Flux.1 offers stronger realism and stylistic diversity.
vs Stable Diffusion 3: SD3 was released by Stability AI shortly before Flux.1 and is widely considered to have underperformed expectations. Flux.1 happened to fill this gap.
Fine-Tuned Models Based on Flux
After Flux.1 went open-source, the community quickly produced a large number of fine-tuned models covering various styles:
- Realistic Photography Style: Simulating film grain and professional studio lighting
- Anime/2D Style: Anime-style LoRAs trained on top of Flux
- Specific Artist Styles: Simulating various painting schools
- Product Rendering: Commercial product photography styles
Civitai and other platforms have accumulated a considerable number of Flux series LoRAs and fine-tuned models that can be downloaded and used directly, working well in combination with ComfyUI.
Who Is It For?
AI Art Enthusiasts: If you have been following the progress of AI art generation, Flux.1 is a model you must understand now.
Local Deployment Users with GPUs: With an NVIDIA graphics card with 12GB or more VRAM and ComfyUI installed, the results of Flux.1 Dev will make you feel it's worth the effort.
Creators Requiring Precise Prompt Adherence: If you don't want to rely on "lottery-style" generation every time and need to realize the images in your head with reasonable accuracy, Flux.1's prompt adherence is more reliable.
Commercial Projects Needing AI Image Generation: Flux.1 Schnell (Apache 2.0) can be used directly for commercial purposes without license issues.
Limitations
Flux.1 is not without flaws. Chinese text rendering remains a weak point; while English is strong, complex layouts can still cause errors. In highly stylized scenarios, Midjourney's artistic feel still has a unique charm. The full versions of Flux.1 Dev and Pro have high VRAM requirements; users without high-end GPUs must rely on online services or quantized versions.
Additionally, for deep SDXL users, their LoRA libraries accumulated over the years cannot be directly used on Flux yet; they need to retrain them or find corresponding Flux versions.
Flux.1 represents an important milestone in open-source text-to-image models. Before it, there was a clear quality gap between open-source models and closed commercial services. With the arrival of Flux.1 Pro, this gap has narrowed significantly. For users who want to run high-quality AI art generation locally, this is currently the model most worth investing time in learning.
