If you’ve spent any time in the design or content creation circles, you’re likely already familiar with the name Midjourney. It’s not just another tool to “try out and forget”; it’s one that has genuinely changed workflows for many people. From the early V4 days to the current V7, we’ve hit plenty of pitfalls and uncovered some patterns along the way. This article organizes those experiences as a reference for those who haven’t yet dived in or are just getting started.
What Exactly Is It?
Midjourney is an AI-powered image generation tool: you input a text description (a prompt), and it generates a corresponding image. On the surface, this sounds simple, but its output quality has consistently remained in the top tier among similar products—especially in artistic style, lighting texture, and composition. Many professional designers admit that the results are “visibly beautiful” to the naked eye.
It is primarily used via Discord, though it has also launched a standalone web version (midjourney.com) with a more intuitive interface, suitable for users unfamiliar with Discord. Since the V7 update, the web experience has become quite complete.
Midjourney is an independent company without big-tech backing. Its founder, David Holz, previously worked on Leap Motion. The company is small (reportedly fewer than 50 full-time employees), yet its product influence is massive—this fact alone speaks volumes.
What V7 Brings to the Table
V7 is the latest major version as of 2025, offering several noticeable upgrades over V6:
More stable character consistency. Previously, generating multiple images of the same person often resulted in drifting facial details. V7 introduces a stronger character reference mechanism; once you set up your reference image correctly, the stability of the same face across different scenes improves significantly. For those working on comics, storyboards, or brand imagery, this is a substantive improvement.
Enhanced detail and realism. Skin texture, fabric folds, and metallic reflections are handled in V7 with a noticeable step up in finesse compared to V6. Especially for realistic portraits, many images look indistinguishable from photographs at first glance.
Improved understanding of long descriptions. With V6, if your prompt was too long or complex, the model would often “selectively ignore” certain elements. V7 offers higher fidelity for complex descriptions, handling the coexistence of multiple elements much more steadily.
How to Write Effective Prompts
This is where many users feel most frustrated at the start: typing a few random words often yields mediocre results or images that bear no resemblance to what you envisioned. Writing prompts requires technique, but it doesn’t need to be overly mystical.
Structured Approach: Subject + Environment + Style + Technical Parameters
For example, if you want an image of a city nightscape:
a lone figure walking on a rainy Tokyo street at night, neon reflections on wet pavement, cinematic lighting, shot on 35mm film, shallow depth of field, photorealistic --ar 16:9 --v 7Let’s break it down:
- Subject:
a lone figure walking on a rainy Tokyo street at night - Lighting/Environment:
neon reflections on wet pavement, cinematic lighting - Style:
shot on 35mm film, shallow depth of field, photorealistic - Parameters:
--ar 16:9(landscape aspect ratio),--v 7(specify version)
Can you use Chinese descriptions? Yes, but the results are generally inferior to English. The model has seen far more English prompts than Chinese ones, so English descriptions typically yield more accurate detail reproduction. If your English isn’t strong, you can use ChatGPT to help optimize your Chinese requirements into an English prompt—a common practice for many users.
Common Parameter Guide:
--ar: Image aspect ratio, e.g.,--ar 16:9(landscape),--ar 9:16(portrait/mobile wallpaper),--ar 1:1(square)--v 7: Specify the latest model version--chaos: Randomness level (0–100). Higher numbers create greater variation between images, useful for exploring styles.--no: Exclude elements, e.g.,--no text, watermark, extra fingersto reduce common errors.--style raw: Reduces AI’s automatic beautification, staying more faithful to your description.--iw: Image reference weight (0.5–2), used in conjunction with reference images.
Who Is It For?
Designers and Illustrators: Quickly generate concept sketches, style references, or visual drafts for client proposals. The efficiency boost is tangible. Instead of spending hours discussing styles with clients, you can generate three or four directions on the spot and let them choose. Brand design, packaging, and poster creation are high-frequency use cases.
Content Creators: For WeChat Official Accounts, Xiaohongshu (Little Red Book), or YouTube thumbnails, if you needabundant illustrations (lots of illustrations) but lack a dedicated designer, Midjourney is the most cost-effective solution.
Game and Film Industry Professionals: Use it during concept design to quickly validate visual directions. Many indie game developers use it directly for character and environment assets, while film producers use it for storyboard references and art concept validation.
Photographers: Useful for pre-visualization, planning shooting scenes and lighting effects. Some also use it to generate composite backgrounds.
Complete Beginners with No Design Background: As long as you’re willing to spend time experimenting with prompts, the barrier to creating good-looking images is not high. However, if you seek highly precise custom results, some learning investment is necessary.
How Does It Compare to Competitors?
vs. DALL-E 3 (Built into ChatGPT): DALL-E 3 excels in Chinese understanding and conversational flow, generating images with text more accurately. However, its artistic feel and stylization fall far short of Midjourney. It is “more correct” but also “more boring,” often failing to satisfy users with high aesthetic standards.
vs. Stable Diffusion: SD is completely free, open-source, and deployable locally, making it ideal for technically skilled users seeking deep customization. However, the learning curve is steep; just setting up the environment can deter many. Output quality heavily depends on the models and workflows you use. Midjourney wins on being ready-to-use with high stability, requiring no technical background.
vs. Adobe Firefly: Firefly has clear commercial copyright licensing (trained entirely on authorized data), making it suitable for commercial scenarios with copyright concerns. It also integrates deeply with Adobe tools like Photoshop. However, in terms of output quality and stylistic diversity, it currently lags behind.
vs. Domestic Products (Jimeng, Tongyi Wanxiang): These have improved rapidly in recent years, leveraging Chinese language understanding as an advantage and offering lower prices. Character generation is decent. However, there is still a gap in overall aesthetic texture and stylization capabilities compared to Midjourney, primarily in that “unclear but visually discernible taste.”
Subscription Pricing
Midjourney does not offer a permanent free plan; it requires a paid subscription (priced in USD):
- Basic Plan: ~$10/month, 200 fast generation hours per month
- Standard Plan: ~$30/month, unlimited slow generations, 15 hours of fast generation per month
- Pro Plan: ~$60/month, 30 hours of fast generation per month, supports Incognito Mode (images do not appear in the public gallery)
- Mega Plan: ~$120/month, 60 hours of fast generation per month, intended for heavy commercial use
Regarding Commercial Copyright: Only Pro plan users and above (or enterprises with annual revenue exceeding $1 million) have full commercial usage rights. For commercial projects, be sure to review the subscription terms carefully.
If you’re just playing around occasionally, the Basic plan is sufficient. If Midjourney serves as a productivity tool for your design work, the Standard plan’s cost is highly justified compared to the labor it saves—outsourced design drafts often cost hundreds or thousands of dollars, while Midjourney handles initial drafts and inspiration references, saving time far exceeding the monthly fee.
Practical Usage Notes
Be Careful with Copyright Issues. Avoid writing specific living artists’ names in your prompts to mimic their style; this is a legal gray area. Instead, describe stylistic features, such as “oil painting style, thick impasto brushstrokes, warm tones.”
Check Generated Images for Details Before Use. Fingers, text, mirror reflections, and distant crowds are common error points. While V7 has significantly improved hand rendering, it’s good practice to zoom in and confirm details before using the image.
Make Good Use of the “Vary (Subtle)” Feature. If an image is mostly good but needs minor adjustments, use the slight variation feature rather than re-running the prompt. This preserves the parts you like while allowing for fine-tuning. “Vary (Region)” allows local modification of specific areas.
Generate Multiple Batches Before Choosing. Each run produces four images; running the same prompt ten times might yield two or three excellent results, with the rest being mediocre. This is normal. Don’t give up just because the first batch wasn’t satisfactory; adjust your description and try again.
Collect Useful Prompts to Build Your Own Library. Good prompts are accumulated over time. Note which word combinations yield results you like so you can reuse them later.
Getting Started Advice
Start with the web version (midjourney.com), as its interface is much more intuitive than Discord.
Don’t aim for “perfection” on your first try. Simply describe a scene you want to see to get a feel for its capabilities. Then, learn from others’ prompts—Midjourney’s public gallery allows you to view other users’ works and their corresponding prompts, which is the fastest way to learn.
Spend an hour experimenting, and most people can start producing images they’re satisfied with. True mastery takes time, but getting started is not difficult.
Midjourney remains the benchmark product in the AI image generation field. While its learning curve isn’t steep, it has depth. Those who take the time to study prompts will see a vast difference in output quality compared to those who just type a few words casually. If your work or creation involves visual content, it is worth learning seriously.
