Defining Success Metrics: Cycle Time, Query Rate, IR Volume, PDUFA Predictability
Eighteen months into your AI program, the CFO asks the question you should have been ready for since the pilot: "Show me the number." Not the anecdote about the writer who finished a Module 2.5 section in a morning. Not the survey where eighty percent of staff said the tool was helpful. The number. A finance executive evaluates an AI investment the way they evaluate any other capital deployment, against a baseline, net of cost, on a metric that ties to either money or risk, and the strategist who walks in with productivity vibes instead of a defended metric will lose the budget to a project that brought receipts. This lesson is the operational-velocity half of that conversation: the throughput and timeline metrics that hold up to a CFO. The next lesson covers the quality metrics that hold up to a CMO. Together they are the measurement spine of your ROI story, and a number that cannot survive a CFO's pushback is not a metric, it is a claim.
Why the CFO Distrusts Your Productivity Number, and Is Right To
Start by understanding why the obvious metric, "writers are forty percent faster," fails the moment it meets a finance brain. First, it has no baseline: forty percent faster than what, measured how, by whom, on which artifacts? An unbaselined percentage is a number with no denominator, and a CFO has seen a hundred of them. Second, it confuses gross task acceleration with realized value: a writer who drafts a section twice as fast has saved time only if that time converts into something, an earlier submission, absorbed volume growth, redeployed capacity, and if it does not convert, the speed is a vanity metric. Third, it ignores the verification tax: the AI draft is faster to produce but the reconciliation work it requires is real labor that the naive "faster" number quietly omits.
The discipline that fixes all three is to define every metric against a documented pre-AI baseline, net of the full cost, and tied to a value pathway. Before you can claim time-to-first-draft fell, you must know what it was, for which artifact, averaged over enough instances to be real rather than anecdotal. Before you can monetize cycle-time compression, you must show the compressed time landed on the critical path and converted to revenue or capacity. And every productivity number must be stated net of the verification labor, the validation overhead, and the tool cost, because a CFO will find those costs whether or not you disclose them, and the strategist who disclosed them first keeps their credibility. The metrics in this lesson are chosen precisely because each one can be baselined, monetized, and defended, which is what separates a board-ready metric from a conference-slide statistic.
Time-to-First-Draft: The Honest Floor
Time-to-first-draft is the most directly attributable AI metric you have, and for exactly that reason it is the floor of your case, not the ceiling. It measures the elapsed working time from "sources assembled and brief defined" to "a complete first draft exists for human review," for a named artifact such as a Module 2.5.4 efficacy section, an ICSR narrative, or a comparability protocol. It is attributable because the AI's contribution to producing a first draft is direct and large, and it is measurable because both endpoints are crisp. This is the number that moves most dramatically in a pilot, and it is the number least likely to be challenged, which is why you lead the operational case with it.
But you must state it honestly, and honesty here is what makes it survive scrutiny. Time-to-first-draft compression is real value only if a first draft was genuinely a bottleneck, which for blank-page-heavy artifacts it often is, and a writer staring at an empty Module 2.5.4 for three days while waiting for inspiration is a real cost the AI genuinely removes. The trap is implying that first-draft speed equals submission speed, when in regulated writing the draft is the start of the work and verification, review cycles, cross-module reconciliation, and QC are the bulk of the timeline. A CFO who hears "we cut drafting time seventy percent" and then sees submission timelines barely move will conclude you measured the wrong thing, and they will be right. State time-to-first-draft as what it is: a large, attributable, early-pipeline gain that is necessary but not sufficient for the downstream metrics that actually carry the ROI.
The right way to present it is alongside the verification tax, as a net figure. If the AI cuts drafting from twenty-four hours to six but adds four hours of claim reconciliation that pure-human drafting did not require in the same concentrated way, the honest net is fourteen hours saved, not eighteen, and the fourteen-hour number is the one that survives. Presenting the gross and the tax together does two things: it preempts the CFO's first objection, and it demonstrates that you understand the workflow well enough to be trusted on the harder metrics. The strategist who volunteers the verification tax on the easy metric earns the benefit of the doubt on the metrics that are harder to verify.
Time-to-Final: Where the Real Money Hides
Time-to-final measures the elapsed time from brief to a fully reviewed, QC-passed, signature-ready artifact, and it is a far more important number than time-to-first-draft because it captures whether the AI actually compressed the part of the timeline that matters. The gap between time-to-first-draft and time-to-final is the verification-and-review interval, and that interval is where AI value is either realized or evaporated. A well-governed AI workflow can compress time-to-final by making the draft cleaner and more consistent, which reduces review iterations and reconciliation effort. A poorly governed one can expand time-to-final, because a draft full of plausible-but-unverified claims and fabricated cross-references generates more review work, not less, and the strategist must be honest about which one their function is producing.
This is the metric that exposes whether your AI program is real or theatrical. A function can show a spectacular time-to-first-draft improvement and a flat or worse time-to-final, and that combination is the signature of a program that has automated the easy part and dumped the consequences downstream onto reviewers. The reviewers absorb the cost invisibly, the headline drafting metric looks great, and the actual artifact takes just as long to finish. Measuring time-to-final is how you catch this, and it is the metric you must improve to make any claim about submission acceleration. A CFO who understands this, and many do, will ask for time-to-final specifically, because they know time-to-first-draft is the metric a vendor demos and time-to-final is the metric a business runs on.
To move time-to-final, the AI value has to extend past drafting into consistency-checking, cross-module reconciliation, and reference QC, which is exactly the integrated-workflow territory of Level 3. This is why time-to-final improvement typically lags time-to-first-draft improvement by a phase: the first-draft gain shows up in month two, the time-to-final gain shows up when the integrated workflows are deployed and validated in the later quarters. Telling the CFO this phasing in advance is a credibility move, because when the time-to-final number is flat in the first quarter you predicted it would be flat, rather than scrambling to explain a disappointing result you did not see coming. The honest phasing story is more persuasive than an optimistic flat one.
Submission Cycle Time and the Critical-Path Discipline
Submission cycle time is the calendar duration of the whole submission build, and it is the metric closest to money, because compressed cycle time can move a launch forward, extend effective patent exclusivity, and accelerate revenue recognition for a named asset. It is also the metric most easily abused, because the temptation is to claim that any task the AI accelerated shortened the cycle, which is false whenever the accelerated task was not on the critical path. Compressing a task with three weeks of slack saves zero calendar days. The critical-path discipline is the single most important analytical move in the entire operational case: you may claim cycle-time compression only for AI acceleration that landed on the binding constraint of the schedule, and you must be able to show which tasks those were.
This discipline is also what makes the metric defensible to a CFO who will, correctly, probe it. When you claim the AI compressed submission cycle time by three weeks, the CFO's question is "which three weeks, on which path, and what was the prior bottleneck?" The strategist who has done the critical-path analysis answers with the specific sequence of activities that constituted the binding constraint and shows how AI acceleration of those specific activities moved the milestone. The strategist who has not done the analysis offers a hand-wave that collapses under the first follow-up. Submission cycle time is where the operational case either becomes a financial argument or remains marketing, and the dividing line is whether you mapped the critical path.
The monetization, once the compression is real, is straightforward and powerful. For a high-value asset, each week of accelerated launch is worth a quantifiable amount of net present value, and at the front end of exclusivity that figure can be large enough that a few weeks of genuine cycle-time compression dwarfs every hours-saved number in the case. This is why the strategist builds up to submission cycle time rather than leading with it: the hours-saved metrics are the floor, time-to-final is the mechanism, and submission cycle time on the critical path is where the dollars actually are. But the dollars are only real if the compression is real, and the compression is only real on the critical path, which is why the discipline is non-negotiable.
Query Rate and IR Volume: The Rework Economy
Query rate and Information Request volume are the metrics that capture rework, and rework is pure waste that a CFO understands instantly. In clinical operations, query rate is the number of data queries generated per unit of data or per site, and a query is a unit of rework: someone has to find it, raise it, route it, answer it, and reconcile it, consuming time across multiple roles. In the submission world, IR volume is the count of Information Requests from the health authority during review, and each IR is a fire drill that consumes senior time on a tight clock, a Day 74 OND request being the canonical example. An AI workflow that improves consistency and catches errors before they ship should reduce both, and a poorly governed one that introduces fabricated cross-references can increase IR volume, which is precisely the inverse risk the validation framework exists to prevent.
What makes query rate and IR volume powerful in front of a CFO is that they convert directly into avoided cost. A reduced query rate is fewer hours of clinical-operations rework per study, which scales across a portfolio into real money. A reduced IR volume is fewer senior-time fire drills during review and, more importantly, fewer opportunities for a review to slip, because a submission that generates few IRs is a submission tracking smoothly toward its goal date. The strategist quantifies these as avoided-cost streams: average hours per query times reduction in query count times loaded cost, and average senior-hours per IR times reduction in IR count plus the schedule-risk value of fewer review disruptions. These are conservative, defensible, and they speak the CFO's native language of waste eliminated.
There is a measurement subtlety that protects your credibility. Query rate and IR volume are influenced by many factors beyond AI, the inherent complexity of the study, the therapeutic area, the reviewer assigned, so you cannot claim every reduction as an AI effect. The honest approach is to attribute conservatively, to control for the obvious confounders where you can, and to present the reduction as consistent with the AI's error-catching mechanism rather than as proof of sole causation. A CFO trusts a strategist who says "this reduction is partly ours and here is the mechanism" far more than one who claims the entire improvement, because the over-claimer is the one whose numbers fall apart under audit and whose next request gets discounted.
PDUFA-Goal-Date Predictability: The Metric the CFO Secretly Wants Most
The PDUFA goal date is the date the FDA commits to act on an application, and for the business it is the anchor around which launch supply, commercial hiring, manufacturing scale-up, and revenue forecasts are all built. Predictability against that date, the degree to which the submission tracks smoothly toward it without the late surprises that threaten it, is a metric the CFO values enormously even though it is the hardest to attribute. A submission that reaches its PDUFA date on a first cycle, without the major-deficiency Information Requests and review disruptions that can force a delay, lets the entire commercial machine commit with confidence, and the cost of a missed or slipped date, idle launch inventory, stranded commercial spend, delayed revenue, is enormous.
The connection between AI governance and PDUFA predictability runs through everything earlier in this lesson. A submission with low IR volume, a smooth review cycle, and consistent, well-verified content is a submission with fewer ways to slip, and an AI workflow that improves consistency and reduces fabricated content is improving the predictability of the date the CFO is planning the launch around. You cannot claim the AI guaranteed the date, the FDA's review is not yours to control, but you can claim, defensibly, that a better-governed submission process reduces the self-inflicted risks to predictability, the avoidable IRs, the inconsistencies that trigger reviewer doubt, the late-discovered errors that force scrambles. Framed as risk reduction rather than guarantee, PDUFA predictability is a metric the CFO will weight heavily.
The strategic value of this metric is that it reframes the entire AI case from cost-saving to risk-reduction, which is a more durable argument with a finance audience than productivity alone. Productivity savings can be competed away or absorbed; predictability of a multi-hundred-million-dollar launch date is a strategic asset that finance protects. The mature operational case therefore stacks the metrics in ascending order of strategic weight: time-to-first-draft as the attributable floor, time-to-final as the mechanism, submission cycle time as the critical-path money, query and IR volume as the rework economy, and PDUFA predictability as the risk-reduction capstone that ties the whole AI program to the number the CFO loses sleep over. Lead with the floor to establish credibility, build to the capstone to win the argument, and never claim a number you cannot baseline, net, and defend.
Building the Baseline Before You Need It
Every metric in this lesson depends on a baseline, and the most common reason an AI ROI case fails is that no one captured the baseline before the AI changed the process. Once the tool is deployed, the pre-AI numbers are gone, reconstructed from memory and anecdote, and a reconstructed baseline is exactly the kind of soft number a CFO discounts. The discipline that saves you is to instrument the function before the rollout: capture time-to-first-draft, time-to-final, submission cycle time, query rate, and IR volume for a representative set of artifacts under the old process, with enough instances to be statistically meaningful rather than a single anecdote. The baseline is the most valuable measurement work you will do, and it has to happen first, which means the strategist who waits until the CFO asks for the number has already lost.
The baseline also has to be honest about variance. Submission cycle times vary enormously by asset complexity, therapeutic area, and team, so a baseline that is a single average hides the spread that a CFO will probe. The defensible baseline captures the distribution, not just the mean, and segments by the factors that drive variance, so that when you claim a post-AI improvement you are comparing like with like rather than crediting the AI for the difference between a simple supplement and a complex original NDA. This segmentation is also what lets you tell the phasing story honestly, showing where AI moved the number and where the variance was always going to dominate, which is the kind of nuance that builds rather than spends credibility.
Finally, the baseline must be governed as carefully as the metrics it underpins, because in a regulated function the measurement itself can be scrutinized. Document how each baseline figure was captured, from which artifacts, over what period, with what definition of each endpoint, so that when you present the improvement you can defend the comparison. The strategist who can produce the baseline methodology, not just the baseline number, is the one whose ROI case survives the CFO's audit and the CRO's scrutiny alike. A metric is only as defensible as the baseline it is measured against, and a baseline is only as defensible as the rigor with which it was captured. Build it first, build it honestly, and build it to be examined.
Key Takeaways
- A CFO distrusts an unbaselined productivity percentage for three good reasons: no denominator, gross-versus-realized confusion, and the omitted verification tax. Every metric must be defined against a documented pre-AI baseline, stated net of verification labor, validation overhead, and tool cost, and tied to a value pathway, because the costs a strategist hides are the ones a CFO finds and then discounts the whole case.
- Time-to-first-draft is the honest floor: highly attributable, large in a pilot, but necessary-not-sufficient, and it must be presented net of the verification tax. Implying first-draft speed equals submission speed is the error that makes a CFO conclude you measured the wrong thing when timelines barely move; volunteering the tax on the easy metric earns trust on the hard ones.
- Time-to-final is where AI value is realized or evaporated, and it exposes theatrical programs that automate drafting and dump unverified claims on reviewers. A spectacular time-to-first-draft with a flat time-to-final is the signature of automating the easy part; the time-to-final gain lags by a phase until integrated workflows are deployed, and predicting that phasing is a credibility move.
- Submission cycle time is the money metric but only on the critical path, and query rate and IR volume are the rework economy that converts directly to avoided cost. Claim cycle-time compression only for AI acceleration on the binding schedule constraint, attribute query and IR reductions conservatively with the mechanism named, and never claim an entire improvement the AI only partly caused.
- PDUFA-goal-date predictability reframes the case from cost-saving to risk-reduction, which is the more durable finance argument, and every metric depends on a baseline captured before rollout. A better-governed submission reduces the self-inflicted risks to the launch date the CFO plans around, and the strategist who instruments the function and documents the baseline methodology first is the one whose ROI case survives audit.
Skill.re