The Six Things AI is Still Reliably Bad At in 2026
The last lesson named six tasks the model does better than you and told you to hand them over. This one is its mirror and its counterweight: six tasks the model is still reliably bad at in 2026, where handing them over ships defects to real users. If the previous lesson was about curing pride, this one is about curing naivety, the equally dangerous belief that because the model nailed the palette and the copy, it can also be trusted with the hierarchy, the novel interaction, and the brand system. It cannot, not yet, and knowing exactly where the competence ends is the difference between a designer who is accelerated by AI and one who is quietly shipping its mistakes. You leave with a "do not ship without a designer" worklist, the guardrail that keeps the speed of delegation from becoming the spread of slop.
Why This Lesson Is the Necessary Counterweight
There is a specific trap that opens up the moment a designer internalizes the previous lesson. Having seen the model genuinely outperform them at ideation, copy, palettes, and the rest, they over-generalize the competence: if it is this good at those, surely it can handle the next thing too. This is exactly how naivety replaces pride as the failure mode. The model's strength on convergent tasks is real, and it creates a halo that makes its weakness on the next category of task invisible, because the output still looks confident and polished. The whole point of pairing these two lessons is to install both boundaries at once: delegate the six it wins, guard the six it loses, and never let competence on the first set buy unearned trust on the second.
The stakes are asymmetric, which is why this lesson matters more than its mirror. Failing to delegate the tasks the model wins costs you time, an annoying but recoverable loss. Wrongly delegating the tasks the model loses costs your users a broken experience and costs you the trust of the people who catch it, a much harder loss to recover. So if you must err, err toward guarding, and use this lesson's worklist as the explicit list of where guarding is non-negotiable. The pattern from last lesson still governs: these six are all low-volume, high-context, no-clear-correct, or must-be-exactly-right, the inverse profile, and that is precisely why the model fumbles them.
One: Hierarchy at Unfamiliar Density
The model handles hierarchy on familiar, sparse layouts because it has seen ten thousand of them, but put it in an unfamiliar high-density situation, a data-dense dashboard, a complex form, a table with twenty columns, a trading interface, and the hierarchy collapses into a flat field where everything competes. This happens because hierarchy under density is not a reproducible pattern; it is a series of context-specific judgment calls about what this particular user needs to see first given this particular data, and the model has no model of that user. It will give every element equal emphasis, lose the scanning path, and produce something that looks orderly and is unusable, because order is not the same as hierarchy. A designer reads the density, decides the one thing the eye should land on first, and builds the visual path to it. The model averages, and the average of a dense screen is a gray wall. Do not ship dense layouts on the model's hierarchy.
Two: Novel Interaction Patterns
Ask the model for a standard pattern (a dropdown, a modal, a tab bar) and it is fine, because those are conventions it has absorbed. Ask it to invent a genuinely new interaction for a problem that does not have an established solution, and it cannot, because invention is the opposite of averaging. The model's entire mechanism pulls toward the most common existing pattern, which is by definition not novel, so when your problem actually needs a new affordance, the model will confidently hand you a conventional one that does not quite fit and will not tell you it settled. Worse, it cannot reason about whether a new interaction is learnable, discoverable, or accessible, because that requires modeling a human encountering it for the first time, which the model cannot do. Genuine interaction-design invention, the hard creative core of the discipline, remains yours. The model is a magnificent librarian of existing patterns and a poor inventor of new ones.
Three: Brand Consistency Across Formats
A model can make one on-brand asset, but ask it to hold a brand consistently across a dozen formats and contexts, a billboard, a favicon, a motion intro, an email header, a dark-mode variant, and it drifts, because brand consistency across formats is a high-context, holistic judgment about what makes the brand recognizably itself when everything else changes. This is latent-space gravity again: each generation slides toward the model's center, and across a dozen formats the cumulative drift produces a set that is individually plausible and collectively off-brand. Recognizing what must stay constant (the proportion, the spacing rhythm, the one color that carries the brand) versus what can flex by format is exactly the holistic brand judgment the model lacks. It sees each asset in isolation; a brand designer sees the system. Brand-system coherence across formats is a guard-it task, which is why the brand-system-as-code work in L3 exists to give the brand a structure the model can be held to rather than trusted to maintain.
The model wins the convergent tasks and loses the ones that need a model of the user, genuine invention, holistic judgment, or exact precision. Competence on the first set buys no trust on the second.
Four: Motion Timing and Choreography
The model can generate motion, but the timing, easing, and choreography that make motion feel right rather than mechanical is a craft of milliseconds and intention that it does not reliably get. Good motion has a point, it directs attention, communicates a relationship, or provides feedback, and its timing is tuned to feel responsive without being frantic, which is a high-context judgment about the specific interaction and the specific feeling intended. The model produces motion that is technically present and emotionally flat, the animation equivalent of the generic average, often too uniform, too long, or applied where no motion was needed. It also cannot reliably reason about motion accessibility, whether the motion respects reduced-motion preferences or risks triggering vestibular discomfort, which is a real-user-modeling problem. Motion craft, the timing and the restraint and the accessibility, stays human, and the motion-accessibility audits in L2 and L3 exist precisely because generated motion cannot be trusted on these dimensions.
Five: Fine Typography
The model gets the type scale right (that is a convergent, high-frequency pattern from last lesson) but fine typography, the kerning of a wordmark, the optical adjustments, the precise line-height for a specific typeface at a specific size, the hanging punctuation, the considered widow and orphan control, is a craft of subtle, expert judgment that it does not reliably execute. Fine typography is low-volume and must-be-exactly-right, the inverse profile, and it depends on a trained eye noticing things most people feel but cannot name. The model produces typography that is acceptable at a glance and wrong to an expert eye, which is fine for a draft and not fine for a brand wordmark or a piece where the typography is the design. The distinction to hold: type scale and basic hierarchy are conceded (last lesson), but fine typographic craft is guarded, because the precision and the trained eye are exactly what the averaging model lacks.
Six: Culturally Specific Imagery
The model is reliably bad at culturally specific imagery, generating representations that are generic, subtly wrong, or stereotyped when a context requires genuine cultural specificity and sensitivity. This happens for two reasons: the training data's biases mean its average representation of a culture skews toward stereotype or toward the dominant culture's view, and getting cultural specificity right requires lived or researched understanding the model does not have. Asked for imagery representing a specific community, festival, or context, it will produce something that looks plausible to someone outside that culture and wrong, sometimes offensively wrong, to someone inside it. This is high-context (it depends on real cultural knowledge), no-clear-correct (it requires judgment and sensitivity), and high-stakes (getting it wrong is not just bad design but can be harmful and reputationally damaging). Culturally specific imagery must be guarded hard, with real cultural input, never shipped on the model's average, which is one of the places its biases are most dangerous and most invisible to a creator outside the culture.
The Artifact: A "Do Not Ship Without a Designer" Worklist
Here is the deliverable, the guardrail counterpart to last lesson's delegation list. It is the explicit, postable list of task types where AI output must not reach a user without a designer's judgment applied. The six above are the universal core, and like the previous worklist you extend it with your domain's specifics. For each entry, write the task and the specific failure to check for, so the guard is actionable rather than a vague "be careful."
- Dense layouts -> check that hierarchy guides the eye to one thing first; reject the gray wall of equal emphasis.
- Novel interactions -> check that the pattern actually fits the problem and is learnable; reject the conventional default forced onto a non-standard need.
- Brand across formats -> check the set for cumulative drift; reject individually-plausible, collectively-off-brand output.
- Motion -> check timing, purpose, restraint, and reduced-motion behavior; reject flat or gratuitous animation.
- Fine typography -> check kerning, optical adjustments, widows and orphans; reject "acceptable at a glance" where the type is the design.
- Culturally specific imagery -> require real cultural input; reject the model's stereotyped or generic average.
The two worklists together, the "I have been outsourced" list and the "do not ship without a designer" list, are the complete map of your relationship with AI as an L1 designer. The first tells you where to delegate so you reclaim time; the second tells you where to guard so you do not ship slop. Pinned side by side, they convert the abstract "AI helps with some things and not others" into two concrete, actionable lists you and your whole team can run, which is the entire practical payoff of understanding where the model's competence begins and ends.
Why the Boundary Moves, and How to Track It
An honest caveat that keeps you from treating this list as permanent: the boundary moves. Some of these six will get better as models improve, and a task that is "reliably bad" in 2026 may become "good enough" later, the same way text-in-images moved from impossible to Ideogram's specialty. So the worklist is not a stone tablet; it is a living document you re-test periodically using the same pattern logic. When a model ships a major update, take one task from your "do not ship" list and re-run your hardest real example through it: has it crossed the line? Usually not yet, but occasionally yes, and you want to be the designer who notices a task became delegable rather than the one still guarding a task the model now handles, or worse, the one who assumed it improved and shipped a failure that is still real. Track the boundary deliberately, in both directions, and your two worklists stay accurate as the frontier advances.
Notice that the deep skill across both lessons is not the lists themselves but the pattern that generates them: high-volume, low-context, convergent, good-enough-tolerant goes to the model; low-volume, high-context, no-clear-correct, must-be-exactly-right stays with you. Memorize the pattern and you can place any task, current or future, named or unnamed, onto the right worklist yourself. The lists are the examples; the pattern is the durable competency, and it is what lets you keep delegating and guarding correctly long after the specific six on each list have shifted.
The Deeper Truth About Where Design Lives
Step back and notice what the six guarded tasks have in common, because it tells you where the discipline actually lives now. Hierarchy under density, novel interaction, brand coherence, motion craft, fine typography, cultural specificity: every one requires either a model of a specific human, genuine invention, holistic judgment across a system, exact expert precision, or real-world knowledge the model does not possess. These are not peripheral skills; they are close to the core of what design has always been beneath the production tasks that AI absorbed. The model took the convergent production and left you the divergent judgment, the human-modeling, the invention, the systemic coherence, the precision, and the cultural truth, which is to say it left you the part that was always the hard, valuable center of the craft. That is not a coincidence; it is the structure of what averaging can and cannot do. The reassuring conclusion of this uncomfortable pair of lessons is that the model offloaded the grunt work and handed you back, in concentrated form, exactly the work that made design worth doing.
The Most Dangerous Property: Confident, Polished Failure
One trait unites every failure on this list and makes it more hazardous than it should be: the model fails confidently, and its failures look polished. A human who does not know how to handle a dense layout produces something visibly rough that signals "not done." The model produces something that looks finished and is wrong, which is far more dangerous, because the visual cue that normally triggers scrutiny is absent. The gray-wall dashboard is evenly spaced and tidy. The forced-fit interaction has clean states. The off-brand asset set is individually beautiful. The flat motion is smooth. The mis-kerned wordmark is crisp. The stereotyped image is well-rendered. In every case the polish actively conceals the defect, which is precisely why a guard based on looks fails and a guard based on the specific behavioral failure is required. This is the same lesson as the very first one in the program, generation versus understanding, arriving now as a warning: do not let the confidence and polish of guarded-category output buy it the trust it has not earned.
The practical consequence is that your guard checks must be behavioral and specific, never aesthetic. "Does this look done?" will pass every one of the six failures. "Does the hierarchy guide the eye to one thing first? Does this interaction actually fit the problem and is it learnable? Does the set hold the brand across all formats? Does the motion have purpose and respect reduced-motion? Is the kerning right where the type is the design? Did a person from this culture validate this image?" will catch them. The worklist's pairing of each task with its specific failure check exists exactly because the generic question is useless against confident, polished failure. Teach your team to ask the specific question, and the polish stops fooling anyone.
Key Takeaways
- This lesson is the counterweight to the last: it cures naivety (trusting the model on tasks it fails) the way the last cured pride. Competence on convergent tasks creates a halo that hides weakness on the next category, and the stakes are asymmetric: wrongly delegating ships defects to users.
- The six the model is still reliably bad at in 2026: hierarchy at unfamiliar density, novel interaction patterns, brand consistency across formats, motion timing and choreography, fine typography, and culturally specific imagery.
- All six share the inverse profile from last lesson: low-volume, high-context, no-clear-correct, or must-be-exactly-right, which is exactly why the averaging model fumbles them.
- Build a "do not ship without a designer" worklist: each guarded task paired with the specific failure to check for, so the guard is actionable. Pinned beside the "I have been outsourced" list, the two form the complete delegate-and-guard map.
- The boundary moves: re-test guarded tasks after major model updates, because some will become delegable (as text-in-images did) and you want to track the line in both directions rather than run on a stale list.
- The durable skill is the pattern, not the lists: it lets you place any task, current or future, onto the right worklist yourself.
- The guarded six share a deep trait: each needs a model of a specific human, genuine invention, holistic judgment, exact precision, or real-world knowledge the model lacks, which is the hard, valuable center of design. AI took the production and handed you back the part that was always the real work.
Skill.re