Hiring the Second Agent Builder: JD, Loop, and Trial Project
Most agent-engineer job descriptions in 2026 read like they were written by a procurement team copying boilerplate. "Build cutting-edge AI agents. 3+ years of LLM experience. Strong Python skills. Familiarity with LangChain, LangGraph, or similar frameworks." Every candidate worth hiring scrolls past these. The ones who apply are the ones who shouldn't. The hiring loop then compounds the problem: a stock take-home that does not predict on-the-job performance, a system-design interview that rewards architecture astronauts over operators, and reference checks that confirm what the interviewers already decided. This lesson is the playbook for hiring the second agent builder on your team: the JD that filters for the right candidate, the four-stage loop calibrated to what actually predicts performance, and the paid trial project that is the single highest-signal artifact in the entire process — and how to design it so it tells you within five days whether to hire.
Why Hiring the Second Builder Is Harder Than the First
The first agent builder on the team — typically the architect themselves — was hired in a context where the team did not yet exist. The first hire was a leap of faith on both sides. The architect was a generalist with relevant background, demonstrated curiosity, and a willingness to learn the platform stack in public.
The second hire is different. By the second hire, the team has an agent in production, a set of operating disciplines, eval infrastructure, an incident history. The second builder needs to walk in and operate inside that context within a few weeks. The hiring bar therefore is not "smart and curious"; it is "smart, curious, and able to operate within the established discipline without slowing the first builder down."
This is harder because the market still treats "agent engineer" as an emerging category. Most candidates who apply are either too senior (will resist the existing discipline) or too junior (will not be effective without long ramp). The candidate who is exactly right — practiced enough to operate, humble enough to fit, energetic enough to elevate — is rare and is being courted by every team in the market.
The cost of a bad second hire
The cost of getting the second hire wrong is asymmetric. A bad first hire is recoverable because the team is still nascent. A bad second hire is much harder to recover from. The team's first agent's reputation depends on the second hire's work. Customers and stakeholders begin distinguishing "the original builder's work" from "the new person's work." Incidents traced to the new hire's changes erode trust in the entire program.
Conservative estimates of the cost of a mis-hire in a senior IC role land between $300,000 and $1,200,000 in lost productivity, remediation, severance, and reputation damage. For an agent engineer specifically, the high end is more accurate because the work touches customer-facing systems and the remediation includes incident response.
The cost of a slow hire
Symmetrically, the cost of taking too long to hire is severe. The first agent builder is alone on the team. They are the eval owner, the prompt owner, the incident response on-call, the platform vendor relationship owner, the cross-functional liaison. They are burning out. Every week the team operates with one builder is a week of accumulating organizational risk and accumulating personal risk for the architect.
The hiring loop should be designed to make a confident decision in three to four weeks from the first conversation, not three to four months. Faster than that is reckless; slower than that is also reckless in a different way.
The job description is not a marketing document. It is a filter. A well-written JD is one that the wrong candidates skip and the right candidate forwards to a friend with the message "this is finally a real one." Every word should be chosen to discourage applicants who will not thrive and to attract those who will.
Writing the JD: Specifics Over Buzzwords
The standard agent-engineer JD reads like every other JD: bullet points of frameworks the candidate should know, generic language about innovation, vague platitudes about culture. This signals a thoughtless team to thoughtful candidates.
The JD that gets the right candidate to apply has six sections, each anchored in specifics:
Section one: what we are doing, in plain English
The opening paragraph is not about the team's mission. It is about the actual work. What agent has been built. What it does. What metrics it has produced. What is in flight next quarter.
Example opening: "We launched a customer support agent in March 2026. It currently drafts responses to about 1,800 Tier 1 tickets per day; reviewers approve 76% with no edits, edit 16%, reject 8%. The team is one engineer (the architect, who is hiring this role) plus a part-time QA reviewer cohort of four humans. We are about to ship the agent's second use case (refund eligibility analysis) and add Retrieval-Augmented Generation against our knowledge base."
The reader who finds this interesting opens the next section. The reader who would find this boring closes the tab. Both outcomes are good filters.
Section two: what you will do in your first 90 days
Concrete, dated. Not "you will help us build agents" — that's the JD anyone could have written. Instead: "In your first 30 days, you'll pair with the architect on the refund-eligibility rollout, learning our eval discipline and incident response patterns. By day 60, you'll own the agent's nightly regression eval and the on-call rotation. By day 90, you'll own at least one new agent end-to-end, from scoping through eval through deployment."
This section signals two things to the right candidate: the team knows what it is doing, and they will be trusted with real ownership quickly. It signals two things to the wrong candidate: this will not be a comfortable learning role, and they cannot float for six months.
Section three: the technical stack you will be working in
Specific. Not "modern AI stack." The actual technologies. The team's choices. The reasons for the choices.
Example: "We use LangGraph for orchestration. Claude Sonnet 4.5 as the primary model with GPT-5 as the backup. Lakera Guard for prompt-injection defense. Datadog for observability. Snowflake for the agent run database. We pinned versions on all of these in March and revisit pins quarterly. We did not choose CrewAI because [reason]; we did not choose AutoGen because [reason]. Those decisions are revisitable but they are decisions we made with eyes open."
The right candidate reads this and has reactions. They have opinions on LangGraph versus CrewAI. They have used Lakera or not. They have specific questions about the Snowflake schema. The wrong candidate reads buzzwords.
Section four: the operating discipline
This is the section that filters out the most candidates and is also the most important. The team's discipline: how the team works, what is non-negotiable, what behaviors are expected.
Example: "We run weekly evals against a versioned 240-case set, growing about 8 cases per week from incidents. Every agent change requires an eval run before deploy; eval regression blocks deploy. Incidents follow our postmortem template (SRE base plus version drift and eval gap fields). The on-call rotation is one week per builder; you will be on call for one week out of every two until we hire the third builder. We do not ship Friday afternoon. We do not deploy untested. We never give the agent more permissions than the minimum required."
The candidate who reads this and thinks "yes, this is how to do this work" applies enthusiastically. The candidate who reads this and thinks "this is too rigid" self-deselects, which is the correct outcome.
Section five: who we are
Brief. Honest. The first builder by name. The reporting line. The size of the broader team. The relationship to the rest of the company.
Naming the existing builder by name does two things. It signals that the candidate will be working alongside a real, named human. It also invites the candidate to look that human up on LinkedIn before applying. Candidates who do this and find a thoughtful builder feel pre-vetted; candidates who feel intimidated self-select out.
Section six: comp and the practical details
Compensation range. Equity range if applicable. Location (remote, hybrid, in-office, with specifics). Time zone overlap requirements. Visa sponsorship if relevant. The travel expectation if any.
Putting comp in the JD does not lose candidates. It saves time for everyone. The candidates who are out-of-range self-deselect; the candidates who are in-range can have substantive conversations from the first call.
What not to include
The JD does not include: a list of every framework the candidate should know (this signals box-checking); company-wide benefits boilerplate (this is for the offer letter, not the JD); aspirational mission language about "transforming the industry"; demand for years-of-experience numbers that the work does not actually require.
The JD especially does not include phrases that have become red flags for senior candidates: "rock star," "ninja," "10x engineer," "passionate about AI," "willingness to wear many hats," "fast-paced environment." Each of these phrases reduces application quality.
The Four-Stage Loop
The hiring loop has four stages over three to four weeks. Each stage has a specific purpose. The loop is calibrated to what actually predicts performance, not what feels rigorous.
Stage one: the working session call (45 minutes)
The first conversation is not a screen. It is a 45-minute working session between the candidate and the architect.
The first 15 minutes: the architect describes the team, the agent in production, the current operating challenges. Specifics. Concrete numbers. The hardest unsolved problem the team is working on right now.
The middle 25 minutes: a problem the team is genuinely facing. The architect describes it. The candidate and architect think through it together, on a shared document or whiteboard. Not "solve this on the spot" but "what would you want to know about this before having an opinion." The conversation reveals how the candidate thinks about uncertainty.
The last five minutes: candidate questions, candidate's process from here, next-step explanation.
Signal from stage one: does the candidate engage substantively, or do they perform? Does their thinking sharpen the architect's, or do they default to received wisdom? Are they comfortable with the team's actual problem, or do they want to talk about the problems they already know how to solve?
Stage two: the paid trial project (five days, asynchronous)
The highest-signal stage of the entire loop. Detailed below — this gets its own section.
Stage three: the trial debrief and team meet (90 minutes)
After the trial, the candidate spends 90 minutes with the team. Not a panel interview. A working debrief plus a team meet.
First 45 minutes: the candidate presents what they built in the trial, the choices they made, the things they learned, the things they would do differently with more time. The architect and one other team member (often the QA reviewer lead, sometimes the engineering manager) ask follow-up questions. Specific. Sharp. About the choices, not the syntax.
Last 45 minutes: the candidate spends time with people they would work with daily. One short conversation with the QA reviewer lead. One short conversation with someone from product or the customer-facing function. The conversations are not formal interviews — they are working introductions.
Signal from stage three: how does the candidate handle being questioned about their own work? Defensive or curious? Do they own the trade-offs they made, or rationalize? How do they show up in conversation with people who are not engineers — the QA reviewer lead, the customer-facing team member?
Stage four: references plus a final architect conversation (60 minutes)
References are mandatory and the architect calls them personally. Three references: one current or recent manager, one peer engineer who has shipped alongside the candidate, one cross-functional partner (PM, designer, customer success — someone who has worked with the candidate but is not engineering).
The reference calls are not formalities. The architect asks specific questions. "Tell me about a time the candidate disagreed with you about a technical decision. How was it resolved?" "Tell me about a time the candidate was wrong about something important. How did they handle it?" "If you could re-hire the candidate, what would you want them to be different?" The vague reference answers ("they're a great engineer," "very thoughtful") tell the architect the reference is performative; the specific ones tell the architect the reference is real.
After references, the architect has a final 60-minute conversation with the candidate. Open agenda. The architect shares any unresolved concerns from references or earlier stages and gives the candidate a chance to address them directly. The candidate shares any remaining questions about the role.
If both conversations end well, the offer goes out within 48 hours.
What this loop deliberately leaves out
The loop does not include: a coding interview testing data structures and algorithms (does not predict agent-engineering performance); a system-design interview where the candidate sketches a generic LLM application architecture (rewards architecture astronauts over operators); a "culture fit" panel (often a proxy for hire-people-like-us bias); a take-home that takes more than five days (signals an exploitative team and screens out candidates with families or other commitments).
The loop is shorter than what many companies run. This is intentional. A loop that is too long signals that the team does not know what it is looking for and is hoping that more interviews will reveal it. A loop that is calibrated to the actual signal is shorter.
The Paid Trial Project
The single highest-leverage artifact in the loop. The trial project is what reveals how the candidate actually performs on the work the team actually does. Designed correctly, it tells the architect almost everything they need to know within five days.
Why paid
An unpaid trial is exploitative and signals that the company does not value the candidate's time. Senior candidates who have other options will simply decline. The trial should be paid at a market rate for the time it asks for — typically $1,500 to $3,000 for a five-day project at evening hours.
The paid framing also clarifies the contract. The candidate is doing real work and the company is acquiring real signal. Both sides have stakes. The conversation about the work is a peer conversation, not an applicant-pleasing-evaluator conversation.
How long
Five evenings of work, spread over a week. Roughly 10 hours total, with the candidate allowed to decide how to split. The cap is critical: a 20-hour trial filters out candidates with families, candidates with current jobs that don't permit the time, and candidates who have other options.
Five days is enough time for a candidate to scope, build, evaluate, and document a meaningful artifact. It is not enough time to ship to production. The trial measures the candidate's ability to execute a small, complete version of the team's real work.
What the trial project should be
The trial project must satisfy four properties to be predictive:
- Representative. The work mirrors what the candidate would do in the role. Not a clever algorithm. Not a contrived puzzle. Something close to the team's actual work, simplified to be doable in five days.
- Self-contained. The candidate can complete the project without access to production data, production systems, or proprietary tools. Public data and standard tooling.
- Open-ended. Many good answers exist. The candidate makes choices and justifies them. The trial reveals their reasoning, not their ability to reach one predetermined answer.
- Reviewable. The deliverable can be examined in 60 to 90 minutes by the team. Code, eval results, a short writeup. No need to run the candidate's submission in production.
Example trial: the support agent ablation
A trial project the architect can use, adaptable to other domains:
"We have built a customer support agent that answers questions about a fictional SaaS product. Here is the agent prompt, the tool definitions, and a small evaluation set of 25 customer queries with reference answers. The agent currently scores 16/25 on this eval — nine of the cases fail. Your job for the next five days: improve the agent's eval score from 16/25 to as high as you can get it, given the constraint that you may change the prompt, add or remove tools, change the model, or change the orchestration — but the eval set and reference answers are fixed. Document your changes, your reasoning, what you tried that did not work, and the trade-offs you made. Bonus if you can identify cases in the eval set that you think have wrong reference answers and explain why."
What this trial measures: practical eval discipline (does the candidate track score before and after each change?); reasoning about prompts and tools (does the candidate make hypothesis-driven changes or guess?); honesty about failure (does the candidate document what did not work?); critical eye (does the candidate notice that some reference answers might be wrong?); communication (is the writeup clear and concise?).
What this trial does not measure but is fine: speed (the project is paced to be possible in 10 hours); pure coding speed (does not matter for the role); arcane LLM trivia (the candidate can look anything up).
Variants by role focus
Adapt the trial to the role's focus:
- If the role leans evaluation: the trial above works directly.
- If the role leans orchestration: provide a working single-agent system and ask the candidate to refactor it into a multi-agent system or to add a new tool while maintaining the eval score.
- If the role leans data and RAG: provide a corpus and a set of questions; ask the candidate to build a retrieval system and tune it; eval is the score on a held-out question set.
- If the role leans operations and observability: provide a corpus of agent run logs from a "production" system and ask the candidate to identify the failure patterns, build alerts, and propose runbook changes.
The team has one trial per role focus; the architect picks the trial matching the role need at the moment of hiring.
What red flags look like in the trial
Specific patterns to watch for:
- Score improved with no documented changes. Either the candidate found a bug in the eval (potentially good — they should have said so) or they overfit to the eval set in a way they cannot articulate (bad — signals they will overfit in production too).
- Long writeup with thin substance. Pages of prose, light on actual changes. Signals a candidate who knows how to perform engineering more than how to do engineering.
- No mention of what did not work. Real engineering involves dead ends. A trial submission that pretends every approach worked is dishonest at best, naive at worst.
- Changes that improve the score but feel wrong. Adding "do not hallucinate" to the prompt. Adding instructions like "always say yes." Hardcoding cases. These signal a candidate who optimizes for the proxy without understanding the underlying objective.
- Tooling explosion. The candidate brings in five new dependencies, builds a custom framework, refactors the orchestration entirely. Sometimes warranted; often signals an engineer who cannot operate within a defined system.
What green flags look like in the trial
- Score improved with documented hypothesis-driven changes. Each change has a hypothesis, a measurable expected effect, and a documented result.
- Identification of issues with the eval set itself. The candidate spots that two reference answers are wrong, explains why, and offers corrected versions.
- Honest documentation of dead ends. "I tried X expecting Y, got Z, abandoned the approach because..." This is the most predictive single signal of senior judgment.
- Constraint awareness. The candidate notes that one of their improvements increases cost per run by 3x, or latency by 2x, and addresses the trade-off explicitly.
- Communication discipline. The writeup is concise, structured, and respects the reviewer's time. Code is clean, comments are sparse but useful, the README is accurate.
Evaluating Candidates Against the Rubric
The team's evaluation rubric should be explicit. The architect writes it before the first candidate enters the loop and uses it consistently across every candidate.
Four rubric dimensions
Score each candidate from 1 to 5 on each:
- Operational discipline. Does the candidate work within constraints, measure before changing, document trade-offs? Trial output is the primary signal.
- Reasoning under uncertainty. When the candidate doesn't know something, do they make it explicit, propose how to learn, and avoid overclaiming? Stage one conversation is the primary signal.
- Cross-functional fit. Can the candidate communicate with non-engineers, take feedback non-defensively, hold a conversation about trade-offs with a product or customer-success counterpart? Stage three team meet is the primary signal.
- Domain pattern recognition. Has the candidate seen the patterns that make agent engineering different from generic backend work — eval discipline, version drift, prompt injection, escalation paths, model trade-offs? Trial and stage three discussion are the signals.
Calibration meetings
After each candidate, the architect and one other team member meet for 30 minutes to score and discuss. Scoring is independent first, then compared. Disagreements are discussed; the team's calibration sharpens with each candidate.
This is true even if the architect is hiring alone — the calibration discussion happens with the hiring manager or with an outside advisor. The point is to externalize the judgment so it does not become "the architect's gut feel."
Closing the Candidate
The right candidate has options. Closing matters.
The offer call
Made by the architect, not by recruiting. The architect tells the candidate why they are the right person, what the team has decided to offer, and what the candidate should expect in the next 24 to 48 hours from recruiting and HR.
The offer is at the top of the range, not the middle. Negotiating from a high anchor is easier than negotiating up from a middle anchor. The team's reputation in the candidate's network depends partly on whether the candidate feels valued by the offer.
The "thinking it over" period
The candidate will have other conversations. The team supports this — does not pressure, does not impose artificial deadlines, does not call the candidate every day. The architect makes themselves available for any follow-up question. The candidate sets the pace.
An honest framing of the deadline: "We would like to give you ten days to make a decision. If you need more time, tell us, and we will accommodate where we can. We are not hiring this role with anyone else in parallel."
The last line is often a lie at other companies. It should not be a lie at this team. If the team is double-tracking candidates, the candidate finds out and trust evaporates.
The decline scenario
If the candidate declines, the architect asks why and listens. The conversation is not a sales attempt. It is information for the next loop. Was the comp wrong? Was the role wrong? Was the timing wrong? Was the team's reputation wrong? Each piece of information improves the next hire.
The team also stays in touch with declined candidates. The right candidate today might be the right candidate in 18 months. Burning the bridge over a decline is amateur.
Onboarding the New Hire
The work is not done at the signed offer. The first 90 days set the trajectory.
The first week
The new builder pairs with the architect on real work, not training material. Day one: walk through the agent in production. Day two: walk through the eval set, the eval runner, the deployment pipeline. Day three: walk through the incident postmortem corpus. Day four: pair on a real change, however small. Day five: retrospective on the first week, adjustment for the second.
The first month
The new builder owns something small and ships it end-to-end. The smallest meaningful piece of work the architect can hand them. They go through the team's full process — eval before deploy, post-deploy monitoring, post-deploy retrospective. The architect reviews but does not redo.
The first quarter
The new builder owns a complete vertical. One agent end-to-end, or one major capability of the existing agent. They are on call for one week of every two. They have run their first incident, written their first postmortem, presented at one all-hands.
If the first 90 days go well, the second 90 days are about scaling — the new builder takes more ownership, the architect is freed for higher-leverage work, and the team is ready to hire the third builder with confidence that the second hire's success can be repeated.
Anti-Patterns to Avoid
The cargo-cult loop
The team copies the FAANG hiring loop because that is what they were trained on. Multiple coding rounds, system design, behavioral, hiring committee. Six weeks elapsed time. The team gets a candidate who passes the loop but cannot operate in agent engineering, because the loop does not test agent engineering.
Fix: design the loop for the work. Discard the components that do not predict performance.
The architect's-clone bias
The architect, looking for someone they can work with, hires the candidate who reminds them of themselves. Same background. Same preferred frameworks. Same blind spots.
Fix: explicit rubric, calibration discussion with another person, and a deliberate look for the candidate who challenges the architect's defaults productively.
The "perfect candidate" wait
The architect interviews twelve candidates over four months. None are "perfect." The architect waits for one who is. The architect burns out. The team falls further behind.
Fix: the rubric. If the architect scores three or more candidates at the threshold over a six-week window and has not hired any, the bar is wrong. Re-calibrate.
The unpaid trial
The team asks candidates to do a 20-hour project unpaid. Senior candidates decline. The pool collapses to candidates with no other options.
Fix: paid, 10 hours, capped, structured. Treat candidates with the respect their time deserves.
The reference call as formality
Recruiting runs the reference checks. The questions are generic. The references confirm everything the team already wanted to hear. The team learns nothing.
Fix: the architect calls references personally. Specific questions. Unresolved concerns surfaced. The reference call is a real diagnostic, not a checkbox.
The post-offer ghost
The offer goes out. The candidate is given a deadline. The team does not check in. The candidate accepts a different offer because that team kept in touch.
Fix: the architect is available during the thinking period. Short check-ins, no pressure. Respect for the decision, with availability for any question.
Key Takeaways
- The second hire is harder than the first. The team has discipline; the new builder must fit it. Bad second hires damage the entire program's reputation; slow second hires burn out the architect.
- The JD is a filter, not a marketing document. Six sections: what we do (specifics, real numbers), what you will do in 90 days, the technical stack with choices and reasons, the operating discipline (the section that filters out the most candidates), who we are, comp.
- Avoid red-flag phrases that reduce application quality: "rock star," "ninja," "passionate about AI," "willingness to wear many hats," generic framework lists.
- Four-stage loop, three to four weeks total: working session call (45 min), paid trial project (5 days, ~10 hrs, $1.5k-$3k), trial debrief plus team meet (90 min), references plus final architect call (60 min).
- The paid trial is the highest-leverage artifact. Five days, ~10 hours, capped. Must be representative, self-contained, open-ended, reviewable in 90 minutes.
- The "support agent ablation" trial: given an existing agent and a 25-case eval, improve the score from 16/25, document changes and reasoning, dead ends, trade-offs. Bonus: identify wrong reference answers.
- Trial red flags: undocumented score improvements (overfitting), thin writeups, no dead ends mentioned, score-gaming changes, tooling explosion. Trial green flags: hypothesis-driven changes, identification of eval set issues, honest documentation, constraint awareness, communication discipline.
- Four rubric dimensions, scored 1-5: operational discipline, reasoning under uncertainty, cross-functional fit, domain pattern recognition. Calibration meetings after each candidate.
- Closing matters. Offer at top of range. No artificial deadlines. The architect available during the thinking period. Decline conversations are information for the next loop.
- Anti-patterns: cargo-cult loop, architect's-clone bias, "perfect candidate" wait, unpaid trial, reference call as formality, post-offer ghost. Each has a fix; the architect plans for each.
Skill.re