โ†
AI Agent Builders & Citizen Developers
Strategic ยท M15 ยท lesson 15 of 32 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
Reporting Cadence and the Quarterly Agent Review
๐Ÿ“–
now learning

Reporting Cadence and the Quarterly Agent Review

15 min

A measurement system is a system only if it runs on a clock. Three honest metrics and a rigorous counterfactual design produce nothing if they live in an analyst's notebook and surface at the QBR for the first time. They produce trust โ€” and budget โ€” when the same dashboard, the same definitions, and the same finance sign-off appear monthly without fail and culminate in a quarterly review that follows a template the CFO has seen before. The Quarterly Business Review for an agent program is not a marketing exercise; it is the meeting where the program's right to exist is renewed for the next quarter. The Agent Architect who walks into a QBR with a coherent template โ€” wins, near-misses, kills, roadmap, with three months of real numbers tied to the methodology document โ€” leaves the meeting with the next budget approved, the next use case green-lit, and a deeper relationship with finance than she had walking in. The architect who walks in with slides assembled the night before leaves with the program on a 60-day watch. This lesson is the reporting cadence, the QBR template that works in 2026, and a sample deck โ€” three months of real numbers โ€” drawn from the running fintech-support example from the previous two lessons.

The Cadence Stack โ€” Daily, Weekly, Monthly, Quarterly

Mature agent programs operate on four nested cadences. Each cadence has a different audience, a different artifact, and a different time horizon. Confusing them โ€” sending a quarterly-grade narrative to a daily audience, or running a daily-grade trace review in a quarterly meeting โ€” wastes time and erodes credibility.

Daily โ€” operational health

The daily review is the agent on-call's morning standup with the runtime data. Cost-per-resolution moved up or down by what percentage. Escalation rate by route. Latency p50, p95, p99. Error rate by category. The artifacts are the platform dashboard (Datadog, Grafana, the agent platform's native observability โ€” LangSmith, Langfuse, Arize Phoenix, Helicone) and the eval pipeline run report. The audience is the agent team itself. The horizon is the next 24 hours.

Weekly โ€” tactical performance

The weekly review is the agent owner's check-in with the operations leader. Aggregate deflection, time-saved, cost-per-resolution for the week. CSAT and reopen trends. Notable incidents and near-misses. Action items rolling over. The artifact is a one-page summary email or a weekly Slack post in the agent program channel. The audience is the agent owner, operations leader, and any operational stakeholders subscribed to the channel. The horizon is the next 1-4 weeks.

Monthly โ€” strategic performance

The monthly review is the formal report to finance and to the program's executive sponsor. Three-metric dashboard with the filters applied, the counterfactual line, and the guardrails. Variance analysis: where the metric moved against forecast and why. Confidence intervals on the impact estimate. The artifact is the dashboard in the finance BI system (Tableau, Looker, Power BI, Sigma) plus a short written commentary. The audience is the finance partner, the executive sponsor, the agent owner, and other senior stakeholders. The horizon is the next 1-3 months and the running quarter.

Quarterly โ€” strategic review and renewal

The quarterly review is the QBR โ€” the formal business review of the agent program. The output is a decision: continue, expand, contract, kill, or pivot. The artifact is the QBR deck, structured to the template in this lesson. The audience is the executive sponsor, the CFO or her delegate, the operations leader, and typically the agent owner plus one or two of the agent team members closest to the work. The horizon is the next quarter and the next 12 months.

The cadence stack is the architect's calendar. The architect who calendars the dailies, weeklies, monthlies, and quarterlies a year in advance โ€” with the audience, artifact, and horizon documented โ€” produces a measurement system that compounds. The architect who treats each report as a fresh ask spends every reporting period assembling rather than analyzing.

The Quarterly Agent Review Template

The QBR template has converged in 2026 on a five-section structure: wins, near-misses, kills, roadmap, and the ask. Each section has a defined purpose and a defined length. The deck is 15-25 slides โ€” long enough to defend, short enough to be read.

Section one โ€” wins (4-6 slides)

The numbered list of agent rollouts in production at quarter-end, each summarized with: name, use case, in-production date, primary metric, primary metric value with confidence interval, CSAT delta, reopen rate, escalation-quality eval score. The methodology reference for each โ€” version and date โ€” at the bottom of the slide.

The win slides answer three questions for the CFO. First, what is the agent doing? Second, is it doing it well? Third, is the methodology trustworthy? The methodology reference is the answer to the third question. If a number is challenged, the architect points to the methodology document, dated and signed, on the company's document system.

The architect resists the temptation to over-claim. Each win is reported with its confidence interval and its guardrail. A win that looks like a $4M saving with a 95% CI of $3M-$5M is reported as that range. The architect who reports $4M point-estimate loses the CFO at the second QBR.

Section two โ€” near-misses (3-5 slides)

Incidents and quality regressions that did not become outages but came close. Brief description, detection, response, action items completed. The near-miss section is where the program demonstrates that it knows where its edges are. A QBR with zero near-misses is suspicious; either the program has detection gaps or the architect is hiding the near-misses.

Common near-misses to surface in 2026:

  • Platform or model version changes that triggered eval regressions caught in shadow before production impact (the n8n LangChain-node tool-arg incident from February 2026 is the canonical industry-wide example).
  • Prompt regressions caught by the eval pipeline before the production deploy.
  • Guardrail triggers that prevented an agent from taking a high-risk action that would have been wrong.
  • CSAT spikes (downward) in a specific segment that were detected and traced to an agent issue before they spread.
  • Cost-per-resolution spikes traced to a model upgrade that increased token usage; remediated by tuning or rolling back.

Each near-miss is reported with what the program learned and what changed. The CFO is reassured by near-misses that surface; she is alarmed by their absence or by hand-wave summaries that lack action items.

Section three โ€” kills (2-3 slides)

Use cases that were tried and stopped. Brief description, why started, why stopped, what was learned, what the costs were. The kill section is the most counterintuitive section to a junior architect; the instinct is to hide kills. The CFO is more reassured by visible kills than by their absence.

A program with zero kills in a year is a program that does not experiment. A program with five kills in a year and a clear pattern of "we kill within 60 days of identifying the workflow is not a fit" is a program that experiments with discipline.

Costs of kills should be reported honestly: engineering hours invested, platform fees incurred, operational time spent. The CFO needs to see that kills are cheap enough to be repeatable. If a kill costs $400K to discover, the program is over-investing before validation; the CFO will push back on the experimentation discipline.

Section four โ€” roadmap (3-5 slides)

The use cases in flight for the next quarter and the next 12 months. Each entry: name, hypothesized impact, methodology design (A/B, pre/post-with-control, synthetic control), pre-launch finance alignment status, target ship date. Color-coded by readiness: green (methodology aligned, design locked, engineering plan in flight), yellow (methodology in progress, design under discussion), red (use case identified but not yet sized).

The roadmap is where the architect signals where the program is going. The CFO scans the roadmap for ambition (is the program expanding?), discipline (is each use case methodologically grounded?), and concentration risk (is the program over-indexed on one bet?).

The roadmap is also where the architect surfaces dependencies on other teams. Finance pre-alignment for a use case planned for Q3 needs the FP&A partner's time in Q2. Surfacing the dependency in the QBR is the architect's mechanism for booking that time.

Section five โ€” the ask (1-2 slides)

What the program needs from the executive sponsor and the CFO in the next quarter. Headcount, budget, executive air cover, partnership with another team, methodology partnership with finance for a new use case. One ask is fine; three is the maximum. More than three signals the architect did not prioritize.

The ask slide is the meeting's purpose for the executive sponsor. The wins, near-misses, kills, and roadmap establish credibility. The ask converts credibility into action. An architect who skips the ask leaves the CFO with no decision to make and the program loses momentum quarter-over-quarter.

Sample QBR Deck โ€” Three Months of Real Numbers

The fintech-support agent from the previous two lessons is now in Q2 2026, three months after the methodology reset. The QBR deck covers February through April 2026. The numbers are representative of a mid-sized fintech with one production support agent and two operations agents in pilot.

Slide 1 โ€” title

"Customer Operations Agent Program โ€” Q1 2026 Quarterly Business Review. Prepared by [Agent Owner]. Methodology version 2.1 (March 2026). Reviewers: CFO delegate, Head of Customer Operations, VP Engineering."

Slide 2 โ€” executive summary (one slide)

  • Customer support agent (Sierra-based) handled 287,000 conversations in Q1 2026. Net deflection 47% (range 44-50%, 95% CI). CSAT 4.2 (vs 4.3 human baseline; pre-agent was 4.1). Reopen 8.1% (vs 7.4% baseline).
  • Cost-per-resolution Q1: $6.10 blended (95% CI $5.80-$6.40), down from $9.40 pre-agent. Methodology: pre/post with three sister business units as control.
  • Operations agents in pilot: AP invoice processing (Stampli + Claude) saved 11 FTE-equivalent hours per week per analyst (95% CI 9-13). Vendor onboarding agent killed in March (see Kills section).
  • One near-miss in Q1 (model version drift, March 12, caught in shadow). No production incidents.
  • Ask: $1.2M Q2 budget approval for two new use cases (collections triage, dispute documentation).

Slides 3-7 โ€” wins

Slide 3 โ€” Customer Support Agent (Sierra). In production since November 2025. Net deflection 47% (44-50%, 95% CI). CSAT 4.2, reopen 8.1%, escalation-quality eval 91%. Pre/post-with-control vs sister BU. Methodology v2.1. Q1 estimated savings: $2.8M (CI $2.2M-$3.4M). Cumulative since inception: $4.1M (CI $3.4M-$4.8M).

Slide 4 โ€” AP Invoice Processing Agent (Stampli + Claude). Pilot since January 2026. Time saved: 11 hrs/wk per AP analyst (CI 9-13). Rework rate 18% (within target). Pilot covers 7 of 14 AP analysts (synthetic control over the other 7). Q1 estimated savings: $340K (CI $280K-$400K). Methodology v2.1.

Slide 5 โ€” Customer Support Agent quality detail. CSAT distribution (90% 4-5, 7% 3, 3% below 3). Reopen distribution. Escalation pattern (top three categories: complex disputes, fraud-suspected accounts, multi-account households). Adversarial eval pass rate 88% (target 85%).

Slide 6 โ€” Cost-per-resolution stack. Five-layer stack visible: labor $3.10, inference $0.42, platform (Sierra) $1.80, integration $0.45, quality program $0.33. Total $6.10. Pre-agent: $9.40 (all labor). Methodology references both the pre-agent baseline calculation (signed Feb 2025) and the v2.1 methodology (signed March 2026).

Slide 7 โ€” Counterfactual chart. Treatment unit (US support center) vs synthetic control (weighted average of 3 sister BUs in EU and APAC) over 18 months. Pre-period RMSE 4.2% (within target). Three placebo tests (each treats one sister BU as if it received the agent); placebo effects within +/- 8% of zero. Treatment effect: -35% cost-per-resolution, p < 0.01.

Slides 8-10 โ€” near-misses

Slide 8 โ€” March 12 model version drift. Claude Sonnet auto-upgraded behind provider endpoint at 02:14 UTC. Shadow eval caught a 6-point regression on the multi-account-household case category within 41 minutes. Production traffic was paused on the affected route for 7 hours while the team validated the new version. No customer impact. Root cause: provider version pinning was nominal but the endpoint's preferred-version pointer had shifted. Action items completed: explicit version pinning in client code, deploy-time validation against the version expected by the eval set, weekly version-drift report.

Slide 9 โ€” March 28 eval regression on disputes category. A prompt change intended to improve fraud detection accidentally weakened dispute-resolution quality. Caught in pre-deploy eval. Eval gap noted: the dispute case set was light on multi-merchant disputes. 5 new cases added to eval set. Methodology v2.1 updated to require category-level eval pass-rate thresholds, not just overall.

Slide 10 โ€” April 7 cost-per-resolution spike. Cost rose to $7.20 for a 3-day period due to a routing change that sent more high-context conversations to the agent. CSAT held. Investigation revealed the routing change was sound but the cost-stack monitoring lagged; the spike was visible only after the period closed. Action item completed: real-time cost alerting at +15% daily threshold. The spike resolved itself once the routing change was tuned.

Slide 11-12 โ€” kills

Slide 11 โ€” Vendor onboarding agent killed March 2026. Pilot ran February-March on Lindy + custom integrations. Goal: automate 75% of new-vendor record creation. Actual: 41% automation rate after eight weeks of tuning. Root cause: the data-quality variance in vendor submissions was higher than the platform could handle without human review on nearly every record, eliminating the time savings. Total cost: $87K (engineering hours + platform fees). Learning documented in the knowledge base. The use case will be revisited if a data-cleansing upstream agent ships.

Slide 12 โ€” IT ticket triage agent paused. Pilot ran January-February. Goal: triage L1 IT tickets to the right downstream queue. Actual: triage accuracy 79% (target 88%). The platform decision (Microsoft Copilot for Service) was correct but the integration with the in-house IT ticketing system was rougher than estimated. Paused pending integration replatform. Cost: $52K. Will revisit in Q3.

Slides 13-17 โ€” roadmap

Slide 13 โ€” Q2 roadmap. Collections triage (green โ€” methodology aligned March 2026, A/B design, target ship May 15). Dispute documentation (green โ€” methodology aligned April 2026, pre/post-with-control on US vs Canada, target ship June 1). Both with hypothesized impact ranges and assumed costs.

Slide 14 โ€” Q3 roadmap. Vendor data cleansing (yellow โ€” use case identified, methodology under discussion with finance and procurement, target ship August). Customer onboarding chat (yellow โ€” design under discussion, target ship September).

Slide 15 โ€” Q4 outlook. Two use cases in red (identified, not yet sized): outbound win-back campaigns, internal knowledge worker copilot. Both will surface in the next QBR with methodology designs.

Slide 16 โ€” Cross-quarter trend chart. Quarterly cost-per-resolution since program inception. Headcount (flat). Volume (up 12%). Cost-per-resolution (down 35%). The chart is the architect's single most CFO-friendly artifact; it shows the program absorbing growth without adding cost.

Slide 17 โ€” Action item completion rate. 27 action items opened in Q1 across postmortems and methodology updates. 24 closed within target dates (89%). 3 open and aging (with explanation). The completion rate is the cultural artifact; the architect signals that the program executes what it commits to.

Slides 18-19 โ€” the ask

Slide 18 โ€” Q2 ask. Approve $1.2M Q2 budget for collections triage and dispute documentation use cases. Breakdown: $480K platform/inference, $400K engineering and integration, $200K quality and eval infrastructure, $120K change management and training.

Slide 19 โ€” Strategic ask. Finance partnership for a Q3 vendor data cleansing methodology by end of May. Operations partnership on the customer onboarding chat use case for Q3 design work by end of June.

Slides 20-22 โ€” appendix

Slide 20 โ€” Methodology versions. v1.0 (Feb 2025, original), v2.0 (Aug 2025, after fintech restatement, see prior lesson), v2.1 (March 2026, current). Each linked to the document.

Slide 21 โ€” Glossary. Deflection rate, FTE-equivalent, cost-per-resolution, escalation-quality eval, parallel trends. One sentence each, with reference to the methodology.

Slide 22 โ€” Acknowledgments. Finance partner, operations leader, engineering lead, the agent team. The architect explicitly thanks the partnership.

Running the QBR Meeting Itself

The deck is necessary; the meeting is where the renewal happens. The meeting in 2026 typically runs 75-90 minutes. Standard structure:

  • Pre-read (sent 3 business days in advance). The deck. The methodology document version reference. The dashboard URL.
  • Executive summary (10 minutes). The architect walks through the executive summary slide. The audience asks clarifying questions.
  • Wins (20 minutes). The architect walks through the wins. The CFO or her delegate validates the methodology references and the confidence intervals. The operations leader confirms the operational reality matches the metrics.
  • Near-misses and kills (20 minutes). The architect walks through what almost broke and what stopped. The audience probes for action item completion and learning loops.
  • Roadmap (15 minutes). The architect walks through the next quarter and 12 months. The audience challenges the priorities and the methodology readiness.
  • Ask and discussion (10 minutes). The architect presents the ask. The CFO and executive sponsor decide on the spot or commit to a decision date.
  • Wrap (5 minutes). Action items captured. Next QBR scheduled. Decisions on the ask documented.

The architect arrives prepared to defend every number. The methodology document is open in another window. Two team members close to the data sit in to handle specific challenges if they come up. The architect does not improvise on numbers; either the number is in the deck (and the architect knows the methodology behind it) or the architect commits to follow up by end of day with the precise figure.

The Mistakes That Kill QBR Credibility

Five recurring mistakes will sink the meeting independent of how good the underlying program is.

The marketing deck

The deck that looks like it was prepared for an investor pitch โ€” heroic claims, oversaturated colors, customer-quote slides. The CFO reads marketing decks elsewhere; she came to QBR for substance. The agent program QBR should look more like an operations review than a sales pitch.

The buried negative

The near-miss section that minimizes the near-miss, or the kill section that does not appear at all. The CFO trusts the architect who surfaces the bad with the good. The architect who buries the negative produces a deck that looks great and a program that gets watched.

The unsupported number

The slide with a savings number and no methodology reference. The CFO will ask. The architect who cannot answer immediately loses the room. Every number references the methodology document; every methodology document is dated, signed, and accessible.

The missing roadmap

The QBR that reports what happened but does not address what is coming. The CFO is approving budget for the future, not for the past. A QBR without a roadmap leaves the audience with no decision to make.

The bloated ask

The QBR with seven asks, none prioritized. The audience cannot decide on seven things in 10 minutes. One to three asks, prioritized, with the top ask clearly the top ask.

The Quarterly Rhythm and the Twelve-Month Arc

QBRs compound over a year. A program with four QBRs that improve in coherence, defensibility, and ambition produces a CFO mental model that maps the program as a reliable investment. By the fourth QBR, the architect has accumulated:

  • A methodology document that has been refined three times based on real review.
  • A roadmap that has shipped 60-80% of what it promised, with the unshipped items either rolled forward with reason or killed with documented learning.
  • An action item completion track record approaching 90%.
  • A library of postmortems and near-miss reports that demonstrate the program's learning loop.
  • A relationship with the finance partner that has survived several real reviews.

The Q5 QBR (the first of the second year) is the renewal point. The CFO either signals the program is a permanent line item or the architect spends the next quarter justifying its existence. The difference between those two outcomes is set by the cumulative quality of QBRs one through four.

Story: The Architect Whose Fifth QBR Was a Formality

A late-stage healthtech company hired an Agent Architect in Q4 2024. Her first QBR (Q1 2025) was attended by the CFO, the CTO, and the Head of Customer Experience. She presented one production agent (intake triage), two near-misses, no kills (the program was four months old), a Q2 roadmap of three use cases, and an ask for $800K. The CFO approved $600K with a request for tightened methodology on the intake triage measurement.

Her Q2 2025 QBR refined the methodology (the diff-in-diff with two sister markets was added), reported two more agents in production, the first kill (an outbound scheduling agent that was killed at week six for quality reasons), four near-misses, and an ask for $1.4M. The CFO approved $1.2M.

Her Q3 2025 QBR introduced the first published cost-per-resolution by use case, a roadmap update, and an ask for headcount. The CFO approved.

Her Q4 2025 QBR introduced the synthetic control analysis for a use case that lacked a natural control. The methodology document was version 2.0. The CFO sent the methodology document to the CEO with a note: "this is what good measurement looks like." The Q1 2026 budget was approved without discussion.

Her Q1 2026 QBR โ€” the fifth โ€” was scheduled for 60 minutes and finished in 35. The CFO had three small questions, the operations leader had two, the CTO confirmed engineering capacity, and the meeting ended. The architect had built a relationship where the QBR was a checkpoint, not a defense. Her Q2 2026 expansion budget was approved before the deck was finalized.

The architect's comment to a peer over coffee: "The fifth QBR is the one where you realize the previous four were the work. The fifth is the harvest."

Key Takeaways

  • The cadence stack has four nested levels. Daily (operational health, agent team), weekly (tactical performance, agent owner + ops), monthly (strategic performance, finance partner + sponsor), quarterly (strategic review, CFO + executives). Each has a defined audience, artifact, and horizon.
  • The QBR template has five sections. Wins (4-6 slides), near-misses (3-5), kills (2-3), roadmap (3-5), and the ask (1-2). 15-25 slides total. Each section has a defined purpose.
  • Wins reference the methodology. Every number cites the methodology document version and date. Confidence intervals, not point estimates. CSAT and reopen guardrails travel with the win.
  • Near-misses and kills earn trust. A QBR with zero near-misses or zero kills is suspicious. A QBR that reports them with action items and learning loops builds the CFO's confidence in the program's discipline.
  • The roadmap is the architect's signaling tool. Green/yellow/red readiness coding. Hypothesized impact and methodology design for each use case. Dependencies on other teams surfaced.
  • The ask is one to three items, prioritized. The CFO came to make a decision; the ask is the decision. Seven asks signal a lack of prioritization.
  • Five mistakes kill QBR credibility. The marketing deck, the buried negative, the unsupported number, the missing roadmap, the bloated ask. Each is preventable; each is common.
  • QBRs compound over a year. The fifth QBR is the renewal point where the program becomes a permanent line item. The cumulative quality of QBRs one through four is what sets that outcome.