Prioritizing Use Cases: The Deloitte/McKinsey Productivity Frame, Critically Examined
Every AI business case in pharma quotes the same numbers. Deloitte's Generative AI in Life Sciences work projects five to seven billion dollars of sector productivity for top-ten biopharma medical writing, and the widely circulated thirty-million-dollars-per-year-per-company savings figure travels with it. McKinsey's "Rewiring Pharma's Regulatory Submissions with AI and Zero-Based Design" puts a twenty to thirty percent reduction in medical-writing effort on the table. These numbers are real, they come from serious firms, and they are also the most misused inputs in the entire field, because they are projections of what is possible under ideal redesign, not guarantees of what your function will capture next quarter. The Level 4 strategist's job is not to repeat these figures to the CRO; it is to understand exactly where the savings come from, where they evaporate, and how a function that prioritizes the wrong use cases can spend a fortune automating a process that should have been eliminated. This lesson takes the productivity frame seriously and takes it apart, so that your use-case prioritization is built on mechanism, not on a number borrowed from a slide.
What the Numbers Actually Say, and Do Not Say
Start with intellectual honesty about the source figures, because a strategist who overstates them loses credibility the first time a CFO asks for the derivation. The Deloitte five-to-seven-billion-dollar projection is a sector-level estimate of productivity potential across the largest biopharma companies, a top-down sizing of an opportunity, not a measured result from a deployed program. The thirty-million-per-company figure is an illustrative allocation of that sector number, useful for framing scale but not a line item any single function has banked and audited. The McKinsey twenty-to-thirty-percent medical-writing-effort reduction is explicitly tied to "zero-based design," meaning it assumes the submission process is redesigned from scratch around AI, not that AI is bolted onto the existing process. That qualifier is the whole story, and it is the part that gets dropped when the number is quoted.
What these figures genuinely establish is that the opportunity is large enough to be worth a serious, funded, multi-year program, and that the direction of the estimate is corroborated across independent serious firms, which matters. What they do not establish is that any specific function will realize twenty to thirty percent by buying a tool, that the savings arrive on a predictable schedule, or that they survive the validation and verification overhead that a regulated function must carry. A strategist who presents the consultancy numbers as the expected return is setting up the function to under-deliver against a borrowed benchmark. The honest framing is that these are credible projections of the ceiling under redesign, and the function's job is to estimate, from its own baseline, how much of that ceiling it can responsibly capture given its readiness and its validation costs. Present the projection as a ceiling and a direction, never as a forecast, and the rest of the business case stays defensible.
Where the Savings Actually Come From
To prioritize use cases, you have to locate the savings in the actual work, and the savings are not evenly distributed across the medical writer's day. The largest, most reliable gains come from the blank-page problem: producing a structured first draft of a document whose shape is highly conventional and well represented in training data. A Module 2.5.4 efficacy section, a Module 2.7.3 efficacy summary, an ICSR case narrative, a standard response document, all have stable structures that an LLM reproduces well, so the first-draft time collapses from hours to minutes. This is real, it is repeatable, and it is where the consultancy numbers are most defensible, because drafting time is a large fraction of writing effort and the structural scaffolding is exactly what the model does best. The second reliable source is consistency and reconciliation at scale: checking a 600-page CSR for internal contradictions, or a Module 2.5 against its 2.7 sub-summaries, work that is tedious and error-prone for humans and pattern-matchable for machines.
The savings come from compressing the production of the conventional and the checking of the consistent. They do not come, and this is the critical asymmetry, from the judgment-dense, low-structure work that actually determines submission quality: the benefit-risk integration in Module 2.5.6, the causality assessment in a complex ICSR, the comparability conclusion under ICH Q5E, the strategic framing of a refuse-to-file response. These tasks are where the senior writer's value concentrates, where the model is least reliable, and where the verification overhead is highest. A use-case prioritization that targets the blank-page and consistency problems captures most of the realizable savings; one that targets the judgment work captures little and risks a great deal, because the verification cost of checking AI-generated judgment can exceed the cost of doing the judgment yourself. The strategist prioritizes where the savings live, which is the high-volume, high-structure, low-judgment work, and explicitly de-prioritizes the inverse.
The Verification Tax That Erodes the Headline
The single most omitted variable in every borrowed productivity number is the verification tax: the time a regulated function must spend reconciling every factual claim and cross-reference in an AI draft to source before it can enter a submission. The consultancy ceiling assumes this tax is small or absorbable; in a regulated function it is neither, and it is the reason the realized savings are always below the projected ceiling. When an LLM drafts a 2.5.4 in minutes, the writer has not saved the document's full production time, because they must now reconcile every hazard ratio, every confidence interval, and every TLF cross-reference against the actual TLF package, exactly the claim-by-claim discipline the foundational lessons teach. That reconciliation is real work, it is non-negotiable, and it must be subtracted from the gross drafting savings to get the net.
This is why the verification tax is the make-or-break variable in use-case prioritization. A use case where verification is cheap relative to drafting, because the output is highly structured and the sources are clean and machine-readable, yields large net savings. A use case where verification is expensive relative to drafting, because the output is judgment-laden and the sources are unstructured, can yield zero or negative net savings even when the gross drafting time collapses impressively. The strategist's prioritization must therefore score each candidate use case on net savings after the verification tax, not gross drafting savings, and the functions that get this wrong are the ones that deploy enthusiastically, measure gross time-to-first-draft, declare victory, and discover at the next budget review that submission cycle time did not actually move because the verification work expanded to fill the gap. The headline number is gross; the defensible number is net, and the difference is the verification tax.
The "Automating the Wrong Process" Trap
The most expensive mistake a function can make is to spend a large AI investment automating a process that should have been eliminated or redesigned, and this is precisely the trap the McKinsey zero-based-design thesis is warning against, though the warning is usually lost when only the percentage survives the retelling. Zero-based design means asking, before automating any step, whether the step should exist at all in an AI-enabled process. A function that has six co-authors manually reconciling a Module 2.5 against its 2.7 sub-summaries through a chain of tracked-changes comments might reflexively deploy AI to accelerate that reconciliation. The zero-based question is sharper: in a redesigned process, would the 2.5 and the 2.7 be generated from a single validated source such that the reconciliation step largely disappears. Automating the reconciliation captures a fraction of the gross savings; eliminating the need for it captures far more, and only the redesign frame surfaces that option.
This trap is why use-case prioritization cannot be a simple ranking of tasks by drafting-time-saved. It must include a redesign lens that asks, for each high-cost process, whether AI should accelerate it, transform it, or eliminate the conditions that make it necessary. The trap is seductive because automating an existing process is easy to scope, easy to fund, and easy to measure, while redesigning a process is hard, political, and slow, so functions default to automation and leave the larger savings on the table. A strategist who understands the consultancy numbers properly knows that the twenty-to-thirty-percent ceiling is a redesign number, which means a function pursuing only automation should expect to capture meaningfully less, and should say so to the CRO rather than promising the redesign ceiling on an automation program. The honest position is that the function will pursue automation for near-term capture and selectively pursue redesign where the savings justify the difficulty, and that the two have different return profiles.
Building the Prioritization Matrix
The output of this lesson is a prioritization matrix that scores each candidate use case on the dimensions that actually predict realized value, not on the borrowed headline. Score four dimensions. First, gross opportunity: how much human effort the task consumes today, because a task that takes ten minutes a year cannot yield meaningful savings no matter how automatable. Second, structure and groundability: how conventional the output is and how clean and machine-readable the sources are, which together predict both how well the model performs and how cheap the verification will be. Third, risk and verification tax: how judgment-laden the task is and how catastrophic an escaped error would be, which together set the verification cost and the deployment caution. Fourth, redesign potential: whether the larger opportunity is to eliminate rather than accelerate the task. A use case that scores high on opportunity and structure and low on verification tax, with no better redesign option, is a first-wave deployment; the inverse is a deliberate non-deployment.
This matrix is what connects the productivity frame to the roadmap from the previous lesson and the business case in the next. The first-wave use cases, high opportunity, high structure, low verification tax, are what phase one of the roadmap deploys, because they capture real savings under heavy supervision without betting on judgment automation. The redesign opportunities are longer-horizon plays that may belong to phase two or three, because they require process transformation that the validation posture must mature to support. And the judgment-dense, high-verification-tax tasks are explicitly documented as out of scope for automation, which is itself a strategic position the CRO needs to hear, because it tells the CRO that the function will not waste investment chasing savings that the verification tax erases. The matrix turns the seductive but dangerous headline number into a defensible, mechanism-grounded prioritization that the function can fund, deploy, and measure honestly. It is the bridge from "the consultancies say twenty to thirty percent" to "here is the fifteen percent we can actually capture, here is where it comes from, and here is what we are deliberately not automating."
Why the Same Percentage Means Different Things by Function
One subtlety that the headline number conceals, and that a Level 4 strategist must surface, is that the realized capture from the same nominal opportunity differs sharply by function and by document type, even within medical writing. A pharmacovigilance function producing two hundred ICSR narratives a month from structured intake fields will capture a high fraction of the gross drafting savings as net savings, because the narratives are short, highly conventional, and grounded in structured data, so the verification tax per artifact is modest and the volume is large. A regulatory writing function producing a handful of Module 2.5 Clinical Overviews a year will capture far less of the nominal percentage, because each document is long, judgment-dense in its integration sections, and grounded in unstructured TLF packages, so the verification tax per artifact is heavy and the volume is small. The same twenty-to-thirty-percent headline applied to both functions overstates the regulatory function's capture and may understate the PV function's, which is why a single sector percentage is a poor planning input for any specific function.
This function-specific variance is why the prioritization must be done bottom-up within each function rather than top-down from a sector figure, and it is also why the highest-volume, most-structured functions tend to be the right place to start a portfolio of AI deployments. The strategist building a multi-function strategy should expect the PV and literature-surveillance use cases to deliver the cleanest early net savings, the clinical-operations monitoring use cases to deliver strong but verification-heavy savings, and the integrated Module 2 regulatory writing use cases to deliver real but smaller and slower-realized savings that depend heavily on data-infrastructure maturity. Communicating this variance to the CRO prevents the common failure of setting a uniform savings target across functions that have structurally different capture rates, and it positions the function leader as someone who understands the mechanism rather than someone repeating a number. The percentage is not a property of AI; it is a property of the document, the data, and the volume, and the strategist who can say which is which is the one whose prioritization the CRO will trust.
Presenting a Critically Balanced Case to the CRO
When the CRO has already read the Deloitte and McKinsey reports, and most have, the strategist who simply repeats the numbers adds nothing and invites the obvious challenge: "then why have we not captured this yet." The strategist who adds value is the one who explains, from the mechanism, why the headline ceiling and the realizable floor differ for this specific function, and who presents a number the function can actually hit. This is the critically balanced posture: take the consultancy projections seriously as evidence that the opportunity is real and large, and take them apart honestly to show why the realized capture is bounded by the verification tax, the readiness gaps, and the redesign difficulty. A CRO trusts the strategist who says "the ceiling is twenty to thirty percent under full redesign, and given our readiness and our verification overhead, our defensible first-wave target is in the low-to-mid teens, growing as the roadmap matures," far more than the one who promises the ceiling.
Frame the balance explicitly so it cannot be mistaken for either hype or skepticism. The case is that the opportunity is genuine and corroborated across serious independent firms, that the mechanism of savings is well understood and concentrated in high-volume high-structure work, that the verification tax and the regulated environment bound the realized capture below the redesign ceiling, and that the largest gains require process redesign the function will pursue selectively rather than universally. This is not hedging; it is the only intellectually honest reading of the evidence, and it is the reading that survives the CFO's request for a derivation and the inspector's question about how the savings were achieved without cutting verification. The consultancy numbers are the beginning of the analysis, not the conclusion, and the strategist who treats them that way builds a business case, the subject of the next lesson, that books savings the function can actually deliver and defend.
Key Takeaways
- The Deloitte five-to-seven-billion-dollar projection and the McKinsey twenty-to-thirty-percent figure are credible ceilings under ideal redesign, not forecasts of what your function will capture. The McKinsey number is explicitly tied to zero-based design, the qualifier that gets dropped when the percentage is quoted. Present them as a ceiling and a direction, never as an expected return.
- The savings concentrate in high-volume, high-structure, low-judgment work, the blank-page problem and consistency checking, not in the judgment-dense tasks that determine submission quality. Benefit-risk integration, complex causality, and comparability conclusions are where the model is least reliable and verification is most expensive, so they capture little and risk much.
- The verification tax is the most omitted variable: the time a regulated function must spend reconciling every claim and cross-reference to source before an AI draft enters a submission. The headline number is gross drafting savings; the defensible number is net after the verification tax, and a use case can show impressive gross savings with zero or negative net savings.
- The "automating the wrong process" trap is spending a large investment accelerating a step that zero-based design would eliminate. Automating an existing reconciliation captures a fraction; generating the 2.5 and 2.7 from a single validated source so the reconciliation disappears captures far more, and only the redesign lens surfaces that option.
- The prioritization matrix scores gross opportunity, structure and groundability, risk and verification tax, and redesign potential, turning a borrowed headline into a mechanism-grounded plan. Present a critically balanced case to the CRO: the opportunity is real and corroborated, but the realized first-wave capture is bounded below the redesign ceiling, and saying so is what survives a CFO's derivation request and an inspector's question.
Skill.re