The Image-Model Field Guide: Firefly, Midjourney, Ideogram, Recraft, Krea, FLUX, Stable Diffusion
The latent-space lesson taught you why image models have fingerprints. This one is the field guide: which model to actually open for four specific briefs that land on a designer's desk every week. Not a leaderboard, because leaderboards are theater that go stale in a quarter, but a casting director's notebook, the kind that says "for this kind of scene, call this actor." By the end you will have a four-brief picker you can use today, and a way of reasoning about a new model that survives the next release cycle.
Why a Field Guide, Not a Ranking
The instinct when faced with seven image models is to ask "which is best." It is the wrong question, and asking it marks you as someone who has not used these tools in anger. There is no best image model any more than there is a best lens or a best actor. There is the right one for the brief, and the brief has requirements: a mood, a composition, whether legible text is involved, whether the output must be vector, and crucially whether you can legally ship it for a paying client. A ranking flattens all of that into one number and is useless the moment your brief differs from whatever the ranking optimized for. A field guide, by contrast, tells you what each model is for, so you can match it to the brief in front of you. The skill this lesson builds is not memorizing today's pecking order; it is learning to read a brief for its real requirements and cast accordingly, which is durable in a way no ranking is.
So we will walk four real briefs, name the model that wins each and why, and along the way build the reasoning that lets you place an eighth model you have never seen the first time you read its capabilities. The four briefs are deliberately chosen to exercise the different requirement axes: editorial illustration (mood), in-product imagery (fit and consistency), brand photography for a launch (rights and polish), and a vector mark exploration (format). Together they cover most of what a working designer actually generates.
The Seven Models, in One Breath Each
Before the briefs, a fast orientation so the names are not abstract. Think of these as the roster you are casting from.
- Midjourney: the highest aesthetic ceiling and the strongest mood. Cinematic, atmospheric, emotionally weighted. No commercial indemnification, and it is the defendant in the consolidated Disney / Universal / Warner Bros. lawsuit, which matters the moment a client is involved.
- Adobe Firefly: trained on Adobe Stock and licensed sources, tuned for clean and predictable commercial output, and the one major model that carries commercial indemnification on paid Creative Cloud plans. Its 2026 Firefly AI Assistant adds agentic, multi-step workflows across Creative Cloud. The safe, defensible choice.
- Ideogram: the typographer. Reliable, correctly-spelled text inside images, plus brand-friendly compositions. The specialist you call whenever words must appear and be right.
- Recraft: vector and raster, brand-anchor-friendly, strong at consistent designed output and raster-to-SVG. The model for icon sets, spot illustration, and anything that has to live in a design system.
- Krea: real-time generation and strong upscaling, good for fast iterative exploration where you want to steer live rather than prompt-and-wait.
- FLUX and Stable Diffusion: open-weights models you can self-host and fine-tune. More neutral and controllable, the choice for sensitive brand work where you want to own the model and restock its library with your own imagery.
Notice that the roster sorts naturally by the requirement each is best at: mood (Midjourney), defensibility (Firefly), text (Ideogram), vector and system fit (Recraft), live iteration (Krea), and ownership and control (FLUX, Stable Diffusion). Hold that mapping and the four briefs almost cast themselves.
Brief One: Editorial Illustration for a Blog Post
The brief: a striking, atmospheric hero illustration for a thought-leadership blog post. No text in the image, no strict brand-color constraint, no client legal exposure because it is your own publication. The dominant requirement is mood and aesthetic impact; you want the image to make someone stop scrolling.
The cast: Midjourney, clearly. This is exactly the neighborhood where its preference-tuning toward beautiful, shareable, emotionally weighted images is the asset rather than the liability. The usual cautions about Midjourney (no indemnification, the lawsuit) are low-stakes here because the image is editorial, on your own property, and not a client deliverable carrying contractual rights warranties. When mood is the deliverable and rights exposure is low, Midjourney is the answer and it is not close. The reasoning move to internalize: identify the dominant requirement (mood), check whether the model's main downside (rights) is actually live for this brief (it is not), and cast.
Brief Two: In-Product Imagery for an Empty State
The brief: a friendly, on-brand illustration for an empty state inside your SaaS product, one of a set of eight empty-state illustrations that must look like they belong together and match the product's existing visual style. Requirements: style consistency across a set, brand fit, and the ability to live in a design system. No legible text, no photographic realism.
The cast: Recraft. This is its home neighborhood, designed-looking illustration with the consistency to hold a style across eight assets and produce them in a brand-coherent way, including vector output that drops cleanly into your system. Midjourney would make each empty state individually gorgeous and collectively incoherent, because its gravity varies the mood from image to image, which is exactly wrong for a set that must feel unified. This brief teaches the consistency axis: when the requirement is "these must look like siblings," you need a model that holds a style, and that points to Recraft over the higher-ceiling-but-more-variable options.
Brief Three: Brand Photography for a Client's Launch
The brief: photographic-style hero imagery for a paying client's product launch, going on their homepage and paid ads. Requirements now include a hard one the first two did not: you must be able to legally ship this for a client, with the rights and indemnification to back the deliverable. Polish and brand fit matter, but the gating requirement is defensibility.
The cast: Adobe Firefly, specifically on a paid Creative Cloud plan for the commercial indemnification. This is the brief where Midjourney's superior aesthetic ceiling is irrelevant, because no amount of beauty offsets handing a client an asset you cannot warrant the rights to, especially with the Disney / Universal / Warner Bros. litigation making the risk concrete and current. Firefly's clean, centered, brand-safe fingerprint is also a fit for commercial launch imagery, and reference anchoring keeps it on the client's brand. The reasoning move: when a requirement is a hard gate (legal defensibility for a client), it overrides the soft preferences (aesthetic ceiling) entirely. You do not trade rights for beauty on a client deliverable. We go deep on the indemnification map in the next chapter; for casting, the rule is simply that client-shipped imagery routes to the indemnified model.
Soft requirements like mood and polish are traded off against each other. Hard requirements like legal defensibility are gates: when one is live, it decides the cast before aesthetics get a vote.
Brief Four: Vector Mark Exploration
The brief: explore a dozen directions for a logo or icon mark, output as clean vectors you can refine in Illustrator or Figma. Requirements: vector output (not raster you have to trace), divergent range, and designed-not-photographic quality. No mood-photography, no legible-text-in-scene, no client-rights gate at the exploration stage.
The cast: Recraft again, for its native vector and raster-to-SVG strength, with Ideogram as a useful second pass if the mark involves a wordmark where letterforms must be exact. Midjourney and Firefly are wrong here not because they lack quality but because they output raster pixels, and a mark exploration that lands as pixels means tracing everything by hand, which defeats the speed the tool was supposed to buy. This brief teaches the format axis, the most commonly ignored requirement: the deliverable's required file type can override every aesthetic consideration, because a beautiful raster is useless when you needed a vector. Always read the brief for its output format before you cast on looks.
The Artifact: A Four-Brief Image-Model Picker
Here is the deliverable, compact enough to memorize and accurate enough to defend. For each brief, the picker names the dominant requirement, the cast, and the one-line reason, so the decision is traceable rather than a vibe.
- Editorial illustration (mood, low rights exposure) -> Midjourney: highest aesthetic ceiling, and its rights downside is not live on your own property.
- In-product imagery, part of a set (style consistency, system fit) -> Recraft: holds one style across many assets and outputs vector that drops into the system.
- Brand photography for a client (legal defensibility is a hard gate) -> Adobe Firefly on a paid plan: commercial indemnification you can warrant to a client.
- Vector mark exploration (output format is the gate) -> Recraft (plus Ideogram for exact wordmarks): native vector and raster-to-SVG, so you refine instead of trace.
The picker is built to be extended. When a fifth brief arrives, do not ask "which is best." Ask the four diagnostic questions the briefs taught: what is the dominant aesthetic requirement (mood, polish, designed-flat), is legible text in the image, is the output format constrained (vector, set-consistency), and is there a hard legal gate (client deliverable). Those four questions route almost any image brief to the right model, and they keep working when the roster changes, because they are about the brief, not the tool.
Placing a Model You Have Never Seen
The real test of understanding is whether you can cast a model that did not exist when you learned the roster. Here is the procedure, and it is the durable skill this whole chapter is building toward. When a new image model launches, ignore the launch-day benchmark theater and ask the requirement questions of it directly. Run your own neutral test prompt and one prompt per requirement axis: a moody scene (does it have aesthetic ceiling?), a scene with a headline (does text render correctly?), a request for a flat icon set (does it hold a designed style and output vector?), and check its terms of service for indemnification (can you ship it for a client?). In twenty minutes you will know which briefs it wins, and you will have placed it in your picker without trusting anyone's leaderboard. This is why the field-guide approach beats the ranking approach: a ranking tells you about the models that existed when it was written, while the requirement-reading skill tells you how to evaluate any model, forever.
Krea, FLUX, and Stable Diffusion slot in through the same lens. Krea wins briefs where the requirement is fast, live, iterative steering, when you want to see the image change as you adjust rather than waiting on full generations. FLUX and Stable Diffusion win briefs where the requirement is ownership and control: sensitive brand work you cannot send to a third party, or a need to fine-tune the model on your own imagery so its home neighborhood becomes your look. None of these are "better" or "worse" than the four-brief winners; they are answers to different requirements, which is the entire point of thinking in briefs rather than rankings.
When a Brief Spans Two Axes: Pipelines, Not Single Casts
Real briefs are not always clean. Sometimes you need both a moody hero and a legible headline on it, or both photographic realism and a vector logo lockup in the same composition. The novice mistake is to hunt for the one model that does both adequately and accept a compromise on each. The better move, and the one a casting director would recognize, is to treat the brief as a pipeline routed by axis rather than a single casting decision. Generate the moody base where mood lives (Midjourney or a fine-tuned open model), then bring the text in where text is reliable (Ideogram, or simply set the type yourself in Figma over the generated image, which is usually the cleanest answer). Generate the photographic scene in the indemnified model for a client, then drop the vector logo in from your actual brand files rather than asking any image model to reproduce a trademark it might mangle.
This reframing matters because it dissolves a lot of false either-or anxiety. You do not have to find a single model that is simultaneously the best cinematographer, typographer, and vector artist, any more than a film needs one person who can act, shoot, and score. You compose specialists. The four-question diagnostic still drives it; you just run the questions and discover the brief has two live requirements, which means two steps, each cast to its specialist. The output is better than any single model would have produced and often faster, because each step plays to a strength instead of fighting a weakness. Thinking in pipelines is the senior move, and it is only visible once you have stopped looking for a single best tool.
The Cost of Casting Wrong
Why does precise casting matter enough to warrant a field guide? Because casting wrong is expensive in three currencies. First, time: forcing Midjourney to produce a consistent eight-illustration set means fighting its gravity through dozens of generations to get coherence Recraft would have given you in the first pass. Second, quality: using a raster model for a vector mark means hand-tracing, and the traced result is worse than a native vector would have been. Third, and most dangerous, risk: shipping a Midjourney image on a client launch because it looked the best exposes the client and your agency to a rights claim that the indemnified Firefly would have prevented. The editorial misjudgment is annoying; the rights misjudgment can end a client relationship or worse. Casting by brief is not pedantry. It is how you avoid paying in all three currencies for a decision that should have taken twenty seconds with the picker in hand.
Key Takeaways
- Do not ask "which image model is best." Ask "which is right for this brief," because the brief has requirements (mood, text, format, rights) that no single ranking can capture and that go stale every release.
- The roster by requirement: Midjourney (mood/aesthetic ceiling, no indemnification), Firefly (clean, defensible, indemnified on paid plans), Ideogram (legible in-image text), Recraft (consistent designed illustration and vector/system fit), Krea (fast live iteration), FLUX/Stable Diffusion (self-hosted ownership and control).
- The four-brief picker: editorial illustration -> Midjourney; in-product set -> Recraft; client brand photography -> Firefly (paid, for indemnification); vector mark -> Recraft (plus Ideogram for wordmarks).
- Hard requirements gate soft ones: legal defensibility and output format override aesthetic preference entirely. You never trade rights for beauty on a client deliverable, and a gorgeous raster is useless when you needed a vector.
- Cast any new model by asking the four diagnostic questions (dominant aesthetic, legible text, output format, legal gate) and running a twenty-minute requirement test, rather than trusting launch-day rankings. This skill survives every release cycle.
- Casting wrong costs in three currencies: time (fighting a model's gravity), quality (tracing a raster you needed as vector), and risk (a rights claim from shipping a non-indemnified image to a client). The picker prevents all three.
Skill.re