Training Data Ethics: The Adobe Firefly 5% Synthetic Problem and Bias in Generated Imagery
"Ethically trained" is the phrase a generative-imagery vendor wants you to repeat, and the phrase a careful designer should never repeat without reading the receipts. In 2024-2025 Adobe disclosed that roughly 5 percent of Adobe Firefly's training set included AI-generated images, some originating from Midjourney - which complicates the clean story that Firefly was built only on licensed and public-domain content. This lesson translates that disclosure into a practitioner-grade understanding of what "ethically trained" actually means, then turns to the bias every image model still ships, with a live audit of how four 2026 models render "doctor," "CEO," and "family." You walk away with an Audit-Before-Use Image Checklist and a set of corrective prompt patterns your brand team can apply the same afternoon.
Why "Ethically Trained" Is a Slogan, Not a Spec
The appeal of "ethically trained" is obvious. A brand team that has to defend its tooling wants a clean sentence to put in a deck, and "we use the ethically trained model" is a clean sentence. The problem is that "ethical" is not a measurable property of a model the way "indemnified" is a measurable property of a contract. It is a marketing claim that compresses a messy, evolving training process into two reassuring words, and the compression is exactly where the truth gets lost.
This is not a reason to be cynical about Adobe specifically. Firefly's licensing posture is genuinely more conservative than most, and its commercial indemnity, covered in the previous lesson, is real and useful. The point is narrower and more useful: a designer who repeats "Firefly is ethically trained" as if it were a fact will get corrected the moment someone in the room has read the actual disclosures, and that correction costs credibility. The professional move is to know the receipts so you can state what is actually true, with its asterisks, rather than parroting the slogan.
The 5 Percent Synthetic Disclosure, Without the Drama
Here are the receipts, stated plainly. Over 2024 and 2025, it became public that approximately 5 percent of the images in Firefly's training set were themselves AI-generated, and that some of those AI-generated images came from other models including Midjourney. Adobe's broader story had been that Firefly was trained on Adobe Stock, openly licensed content, and public-domain material, which is what underwrites the "commercially safe" positioning. The 5 percent figure does not erase that story, but it puts an asterisk on it: a small but non-zero slice of the training data was synthetic, and some of that synthetic data traces back to a model that is itself the subject of major IP litigation.
What does this actually mean for a working designer? Less than the drama and more than nothing. It does not invalidate Firefly's indemnity, which is a contractual promise that stands regardless of training-data philosophy. It does mean the "trained only on licensed human work" narrative is not strictly accurate, so you should not say it. And it means the boundary between "ethically trained" and "trained on outputs of other models" is blurrier than the marketing suggests, which is a useful thing to understand about every vendor's claims, not just Adobe's. The lesson is not "Firefly bad." The lesson is "read the disclosure, state it precisely, and stop treating any vendor's ethics claim as a load-bearing fact."
The 5 percent figure is not a scandal. It is a reminder that "ethically trained" is a story a vendor tells, and your job is to read the footnotes before you repeat the headline.
The Bias Every Model Still Ships
There is a second, larger ethics problem that no indemnity and no licensing posture fixes: every image model still ships demographic bias, because every image model is trained on a corpus that reflects the biases of the internet and the stock libraries it was scraped from. This is the same root cause as the avatar-uniformity tell from earlier in the program, viewed at the level of harm rather than the level of finish. The model averages its corpus, and the average of a biased corpus is a biased image. When the bias is in your brand imagery, it is no longer a craft tell; it is a representation failure that excludes real users and can embarrass the brand.
The "Doctor," "CEO," "Family" Audit
The fastest way to see the bias is to run three loaded prompts across your tools and look at what comes back, unprompted for diversity. Prompt "a doctor" and watch how often the result skews to one gender and one ethnicity. Prompt "a CEO" and watch the skew sharpen - the corpus is saturated with a particular image of executive power, and the model reproduces it. Prompt "a family" and watch how narrow the default family becomes: a particular size, a particular composition, a particular set of assumptions about who counts. Across four 2026 models, the specifics differ but the pattern holds: ungoverned prompts return the statistical center of a biased corpus, and that center under-represents women in authority roles, under-represents non-white people in professional contexts, and under-represents disability, age diversity, and non-traditional family structures almost everywhere.
This is not a hypothetical risk you might encounter. It is the default behavior you will get every time unless you intervene, which is why the audit is a habit and not a one-time exercise. The model is not malicious; it is statistical, and the statistics of its training data are biased, so the output is biased until you correct it.
Why the Bias Is the Designer's Problem, Specifically
It would be convenient to file this under "a research problem for the model makers." It is not, because the designer is the last human between the biased default and the shipped asset. The brand designer generating launch imagery, the graphic designer producing in-product illustrations, the product designer populating an empty state with example avatars - each one is the point where a biased average becomes a published representation that real users will see and read as a statement about who the product is for. If a healthcare app's "your doctor" illustration is always the same demographic, the app has quietly told a large fraction of its users that it did not picture them. The model produced the bias; the designer shipped it. Responsibility lands where the publish button is.
This reframes bias-checking from a compliance chore into a craft and trust obligation. Representative imagery is not a box to tick; it is part of whether the product feels like it was built for its actual users. The designer who audits and corrects is doing the same kind of work as the designer who fixes hierarchy or microcopy: making the artifact actually serve the people who will use it, rather than the statistical ghost of the training corpus.
Corrective Prompt Patterns That Actually Work
The good news is that the correction is mostly within your control at the prompt and review stage. The patterns are simple and they compound. First, specify rather than assume: instead of "a doctor," prompt for the specific, varied people you actually want, and generate a range rather than accepting the first default. Second, request explicit diversity across the dimensions the corpus under-represents - gender, ethnicity, age, ability, body type, and family structure - rather than hoping the model volunteers it, because it will not. Third, generate in sets and curate, so you are choosing from a representative spread rather than shipping whatever the model centered on first.
Fourth, and most important, build the correction into a reusable brand-imagery brief so it is not re-litigated every time. A one-paragraph standard that says "our generated people represent the actual diversity of our users across gender, ethnicity, age, and ability; we generate sets and curate; we never ship the ungoverned default" turns a per-asset judgment call into a team norm. The corrective prompt patterns are the tactical layer; the brief is the strategic one that makes the correction stick. And none of this fights the tool - it directs the tool, which is exactly the relationship a designer should have with a model.
The Audit-Before-Use Image Checklist
Here is the artifact this lesson exists to give you. Run it before any generated human imagery enters a deliverable. It pairs the bias audit with the provenance honesty from the 5 percent discussion, so it covers both "is this representative" and "am I stating its origin truthfully."
- Did I run the bias check? For any prompt involving people, generate a set and look at the spread across gender, ethnicity, age, ability, and family structure. If the set collapses to one archetype, the ungoverned default leaked through.
- Did I correct rather than accept? Apply the corrective prompt patterns - specify, request explicit diversity, generate in sets, curate - rather than shipping the first result.
- Does this representation match my actual users? The standard is not generic diversity; it is whether the imagery reflects the real population the product serves. A statement of who the product is for is being made whether you intend it or not.
- Am I stating provenance honestly? If asked how this was made, can I state the tool and its real claims without overclaiming "ethically trained"? Record the actual posture, asterisks included.
- Is the correction captured in the brief? Has the brand-imagery standard been written down so the next designer does not re-litigate it? A one-time fix that is not codified will regress.
The checklist turns two slippery ethics topics - synthetic-data honesty and demographic bias - into a concrete five-step routine your brand team can run without a philosophy seminar. It is deliberately practical, because the failure mode of ethics in practice is not bad intentions; it is good intentions that never become a repeatable habit.
A Worked Example: Auditing a Healthcare Empty State
The product designer needs an illustration for a "find your doctor" empty state in a healthcare app. They prompt "a friendly doctor" and the first results are, predictably, a narrow archetype. Run the checklist. Bias check: the set collapses to the same demographic across all four tools - the ungoverned default leaked. Correct rather than accept: they re-prompt for a varied set of doctors across gender, ethnicity, and age, generate twelve, and curate four. Match actual users: the app serves a diverse urban population, so the curated set is checked against that reality, not generic diversity. Provenance honesty: generated in Firefly on a paid plan; the team records that and does not claim "perfectly ethical imagery," only "indemnified, bias-audited." Captured in brief: they add a line to the brand-imagery standard so the next empty state starts from the corrected default.
In one short cycle, an asset that would have quietly told a large share of users "we did not picture you" instead represents them, and the correction is now a team norm rather than one designer's good day. That is the difference between ethics as a slogan and ethics as a practice.
Why Synthetic Training Data Is a Feedback Loop Worth Understanding
The 5 percent figure deserves one more pass, because the deeper issue is not the number but the loop it reveals. When a model trains partly on the output of other models, the training corpus stops being a record of human-made imagery and starts including the averages other models already produced. Picture a photocopier copying a photocopy: each pass loses a little fidelity and amplifies whatever distortions were already there. When synthetic data enters a training set, the new model can inherit and concentrate the biases and aesthetic tics of the models that generated that data, including the very avatar-uniformity and compositional sameness this program keeps naming.
This is why the disclosure matters beyond the legal asterisk. A model trained partly on Midjourney outputs may carry forward Midjourney's characteristic look and Midjourney's characteristic biases, one layer removed and harder to trace. The 5 percent is small enough not to dominate Firefly's behavior, but the principle generalizes to every vendor: as the open internet fills with generated imagery, every future model trained on scraped data ingests more synthetic material, and the loop tightens. The practitioner takeaway is not alarm; it is that "trained on the internet" increasingly means "trained partly on other models' averages," so the bias you audit for is not just human bias anymore - it is compounded machine bias, and it is getting harder to attribute to a single source.
For your daily work this changes nothing about the method and everything about the humility. You audit the output regardless of the vendor's training story, because you can no longer assume any corpus is purely human, and you state provenance with the asterisk because the clean "licensed human work only" narrative is becoming structurally less true across the whole field, not just at one company. The loop is the reason the audit is permanent rather than a phase that better models will fix.
A Common Mistake: Auditing Once and Declaring Victory
The most common way teams fail at this is not refusing to audit; it is auditing once, fixing one asset, feeling virtuous, and then regressing on the very next deadline. Bias correction that is not codified is bias correction that evaporates. The designer who lovingly curates a diverse set of doctors for the healthcare empty state on Monday, then ships the ungoverned "a CEO" default on Friday because the launch was on fire, has not built a practice; they have had a good day. The model's default has not changed, the deadline pressure has not changed, and without a written standard the next asset starts from the biased center all over again.
Watch the regression happen in slow motion. The team celebrates the corrected empty state and moves on. Three weeks later a different designer, who never saw that work, generates launch imagery under deadline, accepts the first plausible set because it looks fine and the clock is loud, and ships a hero where every figure of authority is the same demographic. No one is malicious. No one even remembers the earlier audit. The correction lived in one person's good intentions on one Monday, and intentions do not survive a Wednesday. The asset that excludes a large share of users ships, and the brand quietly tells those users it did not picture them, exactly the failure the first audit was supposed to prevent.
The fix is the fifth checklist item, and it is the one teams skip: capture the correction in the brand-imagery brief so the next designer starts from the corrected default, not the model's default. A one-paragraph standard - "our generated people represent the actual diversity of our users across gender, ethnicity, age, and ability; we generate sets and curate; we never ship the ungoverned default" - turns a per-asset act of virtue into a team norm that survives the person, the deadline, and the staffing change. The audit that is not written down is the audit that regresses, and regression is the rule, not the exception, unless the correction outlives the moment.
How to Talk About This With Receipts, Not Moralizing
The tone matters as much as the substance. Moralizing makes designers tune out and makes stakeholders defensive. Receipts do the opposite. Do not say "AI image generators are unethical" - say "every image model ships demographic bias by default, here is the doctor-CEO-family audit that shows it, and here are the corrective patterns that fix it." Do not say "Firefly lied about being ethical" - say "Firefly disclosed about 5 percent synthetic training data including some Midjourney-sourced images, so I describe it as indemnified rather than ethically pure." The credible posture is specific, sourced, and solution-oriented. You are not the ethics police; you are the professional who read the disclosures, ran the audit, and shipped representative work anyway. That is a far stronger position than outrage, and it is the one that actually changes what gets published.
Key Takeaways
- "Ethically trained" is a marketing slogan, not a measurable spec. Read the receipts so you can state what is actually true with its asterisks, rather than parroting the headline and losing credibility.
- Adobe disclosed in 2024-2025 that roughly 5 percent of Firefly's training set was AI-generated, some sourced from Midjourney. This does not invalidate Firefly's indemnity, but it does mean the "licensed human work only" narrative is not strictly accurate, so do not say it.
- Every image model still ships demographic bias, because it averages a biased corpus. This is the avatar-uniformity tell viewed as harm: a representation failure that excludes real users, not just a craft slip.
- The "doctor," "CEO," "family" audit exposes the bias fast: ungoverned prompts return the statistical center of a biased corpus, under-representing women in authority, non-white professionals, disability, age diversity, and non-traditional families.
- The bias is the designer's problem specifically, because the designer is the last human between the biased default and the shipped asset. The model produced the bias; the designer ships it, so responsibility lands at the publish button.
- Corrective prompt patterns work: specify rather than assume, request explicit diversity across under-represented dimensions, generate in sets and curate, and codify the correction in a reusable brand-imagery brief so it does not regress.
- Run the Audit-Before-Use Image Checklist - bias check, correct, match actual users, state provenance honestly, capture in the brief - and talk about it with receipts, not moralizing.
Skill.re