AI for Designers (UX, Product, Brand)
Strategic · M13 · lesson 13 of 26 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
The Design Lead's Hiring Process: Resume Screen, Portfolio Diagnostic, Take-Home Rubric
📖
now learning

The Design Lead's Hiring Process: Resume Screen, Portfolio Diagnostic, Take-Home Rubric

15 min

In 2026, a candidate can generate a flawless-looking portfolio in an afternoon, write a resume that says "leveraged AI to accelerate design velocity by 40%," and walk into your final round having never made an independent judgment call in their career. The old hiring loop - skim the resume, admire the portfolio, run a take-home, vibe-check the final round - was built for a world where producing polished work was itself the signal. That world is gone. Polished work is now free, which means your entire hiring process has to be rebuilt to detect the one thing that is not free and not fakeable: judgment. This lesson designs the full hiring loop for a senior designer in 2026 - a resume screen that catches the filler, a portfolio diagnostic that probes provenance and override decisions instead of outputs, a take-home that demands both an AI-augmented and a hand-drawn artifact, and a final round where the candidate critiques a generated mock against your design system. You walk away with a published hiring rubric, a portfolio-diagnostic question set, and a take-home brief.

The Thing the Old Loop Can No Longer See

The hiring loop you inherited was an instrument calibrated to measure production capability, because for decades production capability was a reliable proxy for everything else: a designer who could produce a beautiful, coherent, accessible flow had necessarily developed the taste, craft, and judgment to do it. The proxy worked because the only way to make the artifact was to have the underlying capability. In 2026 that link is severed. A candidate can now produce the artifact without the capability, because the model produces the artifact. Every stage of the old loop that measured the artifact is now measuring the model's output filtered through the candidate's ability to prompt, which tells you almost nothing about whether they can do the job.

So the redesign principle is singular and it governs every stage: stop measuring outputs, start measuring judgment. Judgment is the capacity the L1 generation-versus-understanding gap names - knowing what the screen is for, catching where the model behaves wrong, deciding what to override and why. It is the one thing the model cannot supply and the candidate cannot fake by prompting, which makes it the only signal worth building a hiring loop around. Every stage below is a different instrument for detecting the same underlying thing: can this person judge, and can they prove it under conditions where the model cannot do the judging for them.

Stage One: The Resume Screen That Catches the Filler

The resume screen's new job is to filter out the candidates who have mistaken tool-operation for capability, and they announce themselves in a predictable dialect. The tell is the AI-as-protagonist sentence: "Leveraged generative AI to produce 40 concepts in one sprint." "Used ChatGPT to synthesize user research." "Drove a 40% velocity increase with AI-augmented workflows." These sentences make the tool the hero and the candidate the operator, and they reveal someone who thinks the value they added was running the tool. A senior designer's resume in 2026 should make the candidate the protagonist and the tool the instrument: "Caught that the generated onboarding flow buried the primary action below the fold and rebuilt the hierarchy." "Overrode the model's research synthesis after verifying a paraphrased quote had inverted the user's actual meaning."

The rubric move is to score resume bullets on a single axis: who is the protagonist, the designer or the tool? A bullet where the designer judges, catches, overrides, decides scores high. A bullet where the candidate "leveraged," "utilized," or "drove velocity with" the tool scores low, because it describes operating a machine, not doing the job. This is not anti-AI; the best resumes show AI used heavily and judged ruthlessly. The filter is specifically for the candidate who cannot distinguish between using a tool and supplying the judgment the tool lacks - which is exactly the candidate the old loop would have advanced on the strength of an impressive-sounding velocity number.

Polished work is now free. Your hiring loop has exactly one job: detect the thing that is not free and cannot be faked by prompting - judgment. Every stage is a different instrument for measuring the same signal.

Stage Two: The Portfolio Diagnostic That Probes Provenance, Not Output

The portfolio review is where the old loop fails hardest, because a portfolio is now a curated collection of outputs that may have been generated, and admiring the output tells you nothing about who made the decisions. The redesign is to stop reviewing the portfolio as a gallery and start using it as a deposition. You are not asking "is this good?" - you can see it is good, anyone's portfolio looks good now. You are asking, of every consequential decision in the work: who or what made this, and why?

The Provenance Questions

For any screen or artifact, the diagnostic question set probes the chain of custody. "Walk me through how this screen came to be - what did you generate, what did you draw, what did you override?" "Show me a place where the AI gave you something plausible and you rejected it. Why?" "This hierarchy is strong - did the model produce it or did you impose it, and how would I know?" "What did you decide not to delegate to the model on this project, and what was the reasoning?" These questions are designed to be unanswerable by someone who merely prompted their way to the artifact, because that person has no override decisions to recount, no rejected outputs to describe, no reasoning to surface. The candidate who did the judging has a rich, specific, unrehearsable answer to every one. The candidate who prompted has fluent-sounding vagueness.

Reading the Answers

The signal is in the specificity and the override stories. A strong candidate will, unprompted, describe a moment of friction with the model - "it kept giving me this centered hero layout and I knew it would push the form below the fold on mobile, so I redrew it" - because real judgment leaves a trail of friction. A weak candidate describes a frictionless collaboration where the AI "really helped speed things up," which is the dialect of someone who accepted the average. Friction with the model is the tell of judgment; frictionlessness is the tell of its absence. The diagnostic is built to surface that distinction, and it does so by refusing to evaluate the output at all and instead interrogating the decisions behind it.

Stage Three: The Take-Home With Two Artifacts

The take-home is where you create conditions the candidate cannot fake, and the design that does this requires two artifacts from the same brief: one AI-augmented and one hand-drawn. Give the candidate a real-shaped problem - a small flow, an empty state, a dense data view - and ask for two deliverables. First, an AI-augmented version with a required provenance log: what they generated, what they overrode, what they kept, and why, documented per decision. Second, a hand-drawn artifact - a sketched flow, an annotated wireframe, a hand-built component - produced without the model, demonstrating the craft substrate underneath the judgment.

The two-artifact structure is the mechanism. The AI-augmented artifact plus its provenance log tests whether they can supervise the model and articulate their overrides; a candidate who only prompted will have a thin, vague, or fabricated log. The hand-drawn artifact tests whether the craft substrate exists at all, because you cannot judge a craft you cannot do, and a candidate whose entire capability is prompting will produce a hand-drawn artifact that is visibly weaker than their AI-augmented one - which is itself the most diagnostic signal in the whole process. When the hand-drawn work is much worse than the generated work, the generated work was the model's, not theirs. The gap between the two artifacts is the measurement.

The brief should be explicit that both artifacts are required and that the provenance log is graded more heavily than the AI-augmented output itself, because the output is the model's and the log is the candidate's. State the time budget honestly (a senior take-home that respects the candidate's time runs to a few hours, not a free weekend of unpaid labor) and tell them exactly what you are evaluating: not whether the output is impressive, but whether the judgment is real and the craft is present.

Stage Four: The Final-Round Critique Against Your Design System

The final round is the live judgment test, and the design is to put the candidate in the exact situation the job consists of: hand them an AI-generated mock and your team's actual design system, and ask them to critique the mock against the system in real time. This is not a portfolio presentation and not a culture-fit chat. It is the job, simulated. You generate a mock with seeded problems - off-token spacing, a competing-primary-CTA error, a destructive action in the safe position, a hallucinated component that does not exist in your system, a missing focus ring - and you watch what the candidate catches, in what order, and how they reason about the fixes.

The rubric for this round scores three things. Catch rate on the high-stakes errors: do they find the destructive-action misplacement and the hallucinated component, or only the cosmetic off-token spacing? A senior who catches only cosmetics is reading the surface, not the behavior. Provenance reasoning: do they spontaneously ask or infer which parts were generated and treat them with appropriate suspicion, or do they trust the mock as if a person made it? System fluency: do they reference your actual tokens and components - "that card-button isn't in your system, it's a hallucinated PrimaryCardButton" - or critique in generic design-school terms? The candidate who runs your Generated-Mock Audit on the spot, names the hallucinated component against your real system, and prioritizes the destructive-action fix over the spacing is demonstrating the exact judgment the entire loop was built to find. The one who admires the polish and suggests a different shade of blue is showing you, live, that they would ship the average.

The Artifact: The Published Hiring Rubric

The deliverable that makes this lesson real is a published, internal hiring rubric - published meaning shared with everyone on the loop so the bar is consistent and the decision is defensible, not a private vibe in the hiring manager's head. The rubric has one organizing principle (measure judgment, not output) and a scored line for each stage. Resume: protagonist score (designer versus tool). Portfolio diagnostic: provenance specificity and override-story richness. Take-home: provenance-log quality and the gap between the hand-drawn and AI-augmented artifacts. Final round: catch rate on high-stakes errors, provenance reasoning, and system fluency.

Publishing the rubric does two things beyond consistency. It makes the hiring decision legible to the candidate's future teammates and to your own manager, so a contested hire can be defended on evidence rather than taste. And it forces your interviewers to align on what judgment looks like, which is itself a clarifying exercise for a team that may not have articulated its own standard since the tools changed. The portfolio-diagnostic question set and the take-home brief are the two operational artifacts that hang off the rubric: the question set scripts the deposition so every interviewer probes provenance the same way, and the brief specifies the two-artifact take-home so every candidate is measured on the same conditions. Together they convert "we'll know good judgment when we see it" into an instrument that reliably sees it.

The Failure Modes This Loop Prevents

Name the hires this loop is built to stop, because they are the expensive ones. The prompt-monkey senior: impressive portfolio, fluent AI vocabulary, no judgment - advances easily through the old loop and ships the average from a senior title, which is far more damaging than a junior shipping it because nobody re-checks a senior. The provenance diagnostic and the two-artifact take-home catch them. The craft-only nostalgic: beautiful hand skills, refuses to engage with AI, cannot supervise a model and so cannot function in an AI-augmented team. The AI-augmented take-home artifact catches them - they produce a weak provenance log because they did not really use the tool. The fluent faker: rehearsed, plausible answers to predictable questions, falls apart under the live final-round critique because you cannot rehearse catching seeded errors in a mock you have never seen. The live round is specifically the un-rehearsable stage.

Notice the symmetry the loop enforces: it rejects both the candidate who can only prompt and the candidate who refuses to prompt, because the job in 2026 requires both halves - the craft substrate and the supervisory judgment. A loop that only valued AI fluency would hire prompt-monkeys; a loop that only valued hand-craft would hire nostalgics who cannot work the modern toolchain. The two-artifact take-home and the live critique are calibrated to require the integration of both, which is the actual senior-designer capability, and which the old single-axis loop could not measure because it was built before the axes split.

Running It Without Cruelty or Theater

A loop this rigorous can curdle into a hazing ritual if you are not careful, and a cruel loop loses the best candidates first because they have options. Three disciplines keep it humane. Respect the candidate's time: the take-home is a few hours and paid if your jurisdiction or ethics demand it, never a free weekend of spec work. Be transparent about what you are measuring: tell candidates upfront that you are evaluating judgment and provenance, not output polish, because the candidates you want will be relieved and the ones you do not will self-select out. And make the final-round critique collaborative rather than adversarial: you are watching how they think, not trying to trip them, so frame it as "let us look at this generated mock together" and let their reasoning breathe.

The goal is a loop that the candidate you want to hire experiences as the most relevant interview they have ever done - finally, an interview that tests the actual job instead of their ability to perform a portfolio walkthrough. When a strong senior finishes your final round and says "that was the first time anyone asked me about the decisions instead of the pixels," you have built the right loop. It detects judgment, it respects the people, and it produces a hire your team can trust with the senior title precisely because they proved, under un-fakeable conditions, that they supply the understanding the model cannot.

Key Takeaways

  • Polished work is now free, so the old hiring loop - which measured production capability as a proxy for everything else - measures nothing. The governing redesign principle for every stage: stop measuring outputs, start measuring judgment, the one thing the model cannot supply and a candidate cannot fake by prompting.
  • Resume screen: score every bullet on protagonist - the designer (caught, overrode, decided) scores high; the tool ("leveraged AI to drive 40% velocity") scores low. The filter is for candidates who mistake tool-operation for capability.
  • Portfolio diagnostic: use the portfolio as a deposition, not a gallery. Probe provenance and override decisions ("show me where the AI gave you something plausible and you rejected it, and why"). Friction with the model is the tell of judgment; frictionlessness is the tell of its absence.
  • Take-home: require two artifacts from one brief - an AI-augmented version with a heavily-graded provenance log, and a hand-drawn artifact produced without the model. The gap between the two is the measurement; when the hand-drawn work is much weaker, the generated work was the model's, not theirs.
  • Final round: hand the candidate an AI-generated mock with seeded errors plus your real design system and watch them critique it live. Score catch rate on high-stakes errors (destructive-action misplacement, hallucinated components), provenance reasoning, and system fluency. This is the un-rehearsable stage.
  • The artifacts are a published hiring rubric (one line per stage, organized around measuring judgment), a portfolio-diagnostic question set (scripts the deposition), and a two-artifact take-home brief. The loop rejects both the prompt-monkey and the craft-only nostalgic, because the 2026 job requires both the craft substrate and the supervisory judgment - and it must be run with respect for time, transparency about what is measured, and a collaborative final round, or the best candidates leave first.