Tracking Productivity, Quality, and Clinical Outcomes
The Five-Metric Definition Sheet from the last lesson is signed and filed, and now comes the part where most measurement programs quietly die: month two. Baselines were captured, the rollout went live, and then nobody pulled the numbers, because the data lives in four systems that do not talk to each other. The EHR knows when notes were signed. The AI vendor's console knows who actually used the scribe. The clearinghouse knows which claims Anthem denied and why. The MBC platform knows which PHQ-9s came back. Jordan's leadership meeting needs all four on one page, every month, with a clear answer to the question almost nobody asks precisely: is this number a productivity gain, a quality gain, or an outcome gain? Conflating them is how a practice ends up telling a board that an AI scribe improved depression remission, a claim that dissolves under the first skeptical question. This lesson teaches you to build the monthly AI impact dashboard for a therapy group practice, pulling EHR data, AI tool usage data, payer claims data, and MBC outcome data into one disciplined page. By the end you will have written the Monthly Dashboard Spec: sources, refresh cadence, the three-lane gain taxonomy, and the reading protocol for the meeting where the numbers land.
Four Systems, One Page: The Data Plumbing Problem
Start with an inventory of where each metric actually lives, because the dashboard is only as good as its plumbing. Documentation time per session comes from the EHR's timestamp pair, session end to note signature, or from the structured clinician log used at baseline if the EHR cannot produce it. Claim denial rate comes from the billing system or clearinghouse, exported by denial reason code and CPT code. MBC completion and the outcome trajectory come from the measurement platform, Blueprint Health, Greenspace, or Owl, cross-checked against what was filed in the EHR chart, because a score that lives in the platform but never reached the chart is a completion failure. Retention and the quarterly pulse come from payroll and the survey tool. And the fourth source, the one practices forget until the numbers stop making sense, is the AI vendor's own usage data: which clinicians generated drafts, how many, for what session types.
The usage data is not vanity telemetry; it is the denominator that makes every other number interpretable. If documentation time fell 40 percent practice-wide but the console shows only 12 of 25 clinicians using the scribe, the adopters' real improvement is far larger than 40 percent and the program's headroom is the other 13 clinicians, a change-management finding, not a tooling finding. Conversely, if usage is 95 percent and documentation time barely moved, the tool is not delivering and the pilot data needs re-examination before renewal. A dashboard without the usage column cannot distinguish "the tool does not work" from "half the practice is not using the tool," and those diagnoses have opposite remedies and opposite renewal decisions.
One rule governs the plumbing before any chart gets drawn: the dashboard carries practice-level, de-identified data only. No client names, no client-level score lists, no note excerpts. Aggregates, counts, rates, and averages travel to the leadership page; protected health information does not. The dashboard is a business document that will be emailed, screen-shared, and eventually shown to a board member, a banker, a buyer, and it must be built from day one so that nothing on it ever needed a release.
The Three-Lane Taxonomy: Productivity, Quality, Outcome
Here is the controlling analogy, worth holding through every section: the dashboard is the practice's monthly lab panel. A lab panel does not give one health score; it groups markers into panels, metabolic, lipid, hepatic, because markers in different panels mean different things, move on different timescales, and respond to different interventions. Reading a lipid value as a liver enzyme is not aggressive interpretation; it is a category error. The dashboard's three panels are productivity, quality, and outcome, and every number on the page belongs to exactly one.
Productivity gains are time and throughput: documentation minutes, signature latency, after-hours charting, usage rates. They move fast, within weeks, and they are what an AI scribe most directly produces. Quality gains are about the work product: denial rate by reason code, the share of notes meeting the documentation standard, MBC completion. They move on a one-to-two-quarter lag, because a better note today becomes a paid claim sixty to ninety days from now, and a measurement cadence takes a quarter to become a habit. Outcome gains are about clients: the aggregated PHQ-9, GAD-7, and PCL-5 trajectories by cohort. They move slowest, on quarters to years, and the AI program's defensible relationship to them is measurability and safety monitoring, never causation. The scribe did not treat the depression; it made the practice able to see whether the depression is being treated, and to detect if the rollout is making care worse.
The taxonomy earns its keep the first time someone tries to move a number between lanes. A vendor renewal deck will happily put "improved client outcomes" next to its time-savings chart; the three-lane dashboard makes that move visibly illegitimate, because the outcome panel carries its own caveat and evidentiary standard. The same discipline protects you in the other direction: when a board member sees a flat outcome panel and asks whether the program failed, the lane labels let you answer that outcomes were never the program's causal claim; the productivity and quality panels are where its promises live, and the outcome panel exists to prove the program did no harm.
Building the Productivity Panel
The productivity panel carries four rows, each monthly, each against its baseline. Row one: documentation time per session, by note type, exactly as the definition sheet specifies, session end to signature. Row two: signature latency, the share of notes signed within 24 hours, the audit-readiness number that pairs with row one so neither can be gamed alone. Row three: after-hours documentation, the share of notes signed after 7 PM or on weekends, the burnout proxy that shows whether the program gave clinicians their evenings back or just compressed daytime charting. Row four: AI usage, drafts per clinician per week from the vendor console, the denominator row.
Work one month of Jordan's numbers to see how the panel reads. Baseline: 21 minutes average per note, 61 percent signed within 24 hours, 34 percent of notes signed after 7 PM, usage zero. Month four: 11 minutes average, 88 percent within 24 hours, 14 percent after 7 PM, and the usage row shows 19 of 25 clinicians generating drafts on at least 80 percent of sessions. The translation, which the panel footer computes automatically: 10 minutes saved per note, across roughly 110 completed sessions per week per the EHR, is about 18 clinician-hours per week returned to the practice. Note what the panel does not say: that those 18 hours became billable hours, revenue, or retention. Converting saved time into dollars is the next lesson's job, and a dashboard that converts prematurely, with optimistic assumptions buried in a cell formula, is how vendor math sneaks into your own reporting.
The panel also carries its honesty marks. Where the timestamp method changed mid-stream, the panel says so. Where six clinicians remain non-users, the average is shown both blended and adopters-only, labeled, because the blended number understates the tool's effect, the adopters-only number overstates the program's, and a careful reader deserves both.
Building the Quality Panel
The quality panel carries three rows on a quarterly comparison cadence, even though the dashboard refreshes monthly, because quality numbers are noisy at monthly resolution and the definition sheet committed to quarterly judgment. Row one: denial rate by reason code and CPT code, with the AI-addressable categories, medical-necessity and documentation-insufficiency denials, separated from the categories AI cannot touch, eligibility and coordination-of-benefits. The addressable categories are the program's quality claim; the untouchable categories are the control series. If both fall together, something other than documentation changed (a payer policy, a biller's process fix) and the program should not claim the movement. That built-in control is what separates a dashboard from a sales chart.
Row two: documentation standard compliance, the share of audited notes meeting the practice's checklist of verifiable details a payer reviewer looks for and AI cannot supply on its own: time-in-session minutes supporting a 90837, the specific PHQ-9 score and delta, the modality actually named rather than "processed feelings." This row is fed by a monthly chart audit, ten notes pulled at random by the clinical director, and it catches the failure mode the productivity panel cannot see: fast, signed, thin notes. A scribe produces fluent boilerplate at speed; only a human audit confirms the session-specific facts that survive a high-utilization 90837 review, and the memory of a colleague's $14,200 recoupment letter is why this row exists.
Row three: MBC completion rate, monthly, against the committed cadence: PHQ-9 every session for the depression caseload, GAD-7 every fourth session for anxiety, PCL-5 monthly for PTSD work. Completion means scored and filed in the chart, and the row shows the trend from baseline, because the journey from a 40 percent paper-form rate toward 85 percent automated completion is the clearest quality story the program will tell, and it compounds: every point of improvement feeds the outcome panel below and the 96127 billing line where coverage applies.
A dashboard is the practice's monthly lab panel: productivity, quality, and outcome are different panels, on different timescales, with different evidentiary standards. Reading one as another is not optimism; it is a category error, and sophisticated readers catch it in seconds.
Building the Outcome Panel, and Its Caveat Line
The outcome panel is the smallest and the most carefully worded. It carries the aggregated trajectory per diagnosis cohort: average PHQ-9 change from baseline to most recent administration for the depression cohort, the GAD-7 equivalent for anxiety, the PCL-5 equivalent for trauma work, each with its response rate printed beside it so a reader can see how much of the caseload the average represents. Clients enter a cohort at two administrations; terminated clients carry their last observation into the termination quarter and exit. The panel reports every cohort every quarter, flat ones included, because the pre-commitment against cherry-picking was made on the definition sheet and the dashboard is where that promise is kept.
The caveat line is printed on the panel itself, not in an appendix nobody reads: "Outcome trajectories are reported for clinical monitoring and program safety. The AI program's claim is measurability, not causation; no inference from documentation tooling to clinical outcomes is made or implied." That sentence does three jobs: it inoculates against the board member who wants "AI improved remission" in an investor update, it protects clinicians from having their outcomes attributed to software, and it preserves the panel's real function, the tripwire. The operational rule: a cohort deteriorating across two consecutive quarters post-rollout triggers a clinical review, led by the clinical director, asking whether anything in the program, rushed sessions, verification fatigue, documentation crowding out attention, is implicated, with rollout expansion paused until the review reports.
Hold the hard boundary here too, stated on the panel spec: the AI layer administers, transports, scores arithmetic, and files. It never interprets a score into a risk level, never scores a CSSRS, never flags a client as "high risk" by its own judgment. The dashboard may show that scores were filed late or completion fell; it never shows an AI-generated risk stratification, because risk determination belongs to the clinician whose license sits behind the chart.
Cadence, Ownership, and the Reading Protocol
A dashboard nobody is accountable for refreshing becomes archaeology by month three, so the spec assigns the owners the definition sheet named: the practice manager pulls the productivity panel from the EHR and vendor console by the 5th business day of each month; the billing manager refreshes the denial rows from the clearinghouse on the same schedule, with the quarterly comparison computed at quarter close; the clinical director runs the ten-note audit and the MBC and outcome rows. One person, not a committee, assembles the page, and the page is one page: if it does not fit, the dashboard has forgotten the lesson about twenty metrics.
Then write the reading protocol, because how numbers are read determines whether the dashboard manages the program or terrorizes the staff. The monthly leadership reading is fifteen minutes, three questions per panel: what moved, what explains it, what decision does it change. Numbers are read at practice level; no clinician is named in the leadership meeting for a productivity or quality number, ever, because the first time the dashboard singles someone out, clinicians start managing the numbers instead of the clients, and the data corrupts exactly the way clinician-level outcome rankings corrupt. Individual coaching happens in supervision, fed by the same data, under that relationship's protections, not on the leadership screen.
Finally, the spec defines escalation thresholds in advance: signature latency below 80 percent within-24-hours triggers a workflow review; an addressable-denial uptick sustained for a quarter triggers an expanded documentation audit; usage below 70 percent at month six triggers the change-management playbook rather than a tooling debate; and the two-quarter outcome deterioration rule triggers the clinical review described above. Pre-committed thresholds prevent the two failure modes of dashboard meetings: panic over noise and rationalization of signal.
Writing the Monthly Narrative: One Paragraph per Lane
Numbers without narrative get misread, so the spec ends each monthly page with a three-paragraph narrative, one per lane, written by the panel's owner in plain declarative sentences. Productivity states the time returned: "Average documentation time held at 11 minutes; after-hours notes fell to 14 percent; 18 clinician-hours per week returned versus baseline; six non-users scheduled for the February training cohort." Quality states the claim and the control: "Medical-necessity denials on 90837 fell from 9.1 to 5.8 percent quarter over quarter while eligibility denials held flat, supporting attribution to documentation content; the ten-note audit passed 9 of 10." Outcome states the monitoring result in the caveat's language: "All three cohorts stable or improving; response rates 78 to 84 percent; no safety threshold triggered."
Notice what the grammar enforces. Productivity sentences use time units. Quality sentences use rates with controls. Outcome sentences use monitoring language, stable, improving, deteriorating, threshold, never causal language. A reader who sees three months of these paragraphs absorbs the taxonomy without being lectured, and a writer forced into this grammar cannot accidentally claim that a scribe treated PTSD. The narrative is also where honest complications live: the month an EHR upgrade broke the timestamp report, the quarter a payer changed its prior-auth policy and the control series moved. Complications reported promptly are credibility; complications discovered later by a board member are a credibility funeral.
This discipline is also the bridge to the next lesson. The ROI memo for owners, boards, and investors is built almost entirely from the productivity and quality paragraphs, translated into dollars with conservative, stated assumptions. A year of well-written narratives makes the memo a compilation exercise; a year of bare numbers makes it forensic reconstruction.
The Applied Problem: Write the Monthly Dashboard Spec
Your artifact is the Monthly Dashboard Spec: a three-to-four-page document that lets anyone in the practice build and refresh the dashboard without you in the room. Draft the skeleton with an AI assistant under your approved-tool policy, using only de-identified, practice-level content. A working prompt: "Create a monthly AI impact dashboard specification for a 25-clinician behavioral health group practice. Structure: (1) data source inventory with four sources (EHR, AI vendor usage console, clearinghouse claims, MBC platform), each with the named report, owner, and 5th-business-day refresh deadline; (2) three panels, productivity, quality, outcome, with the exact rows I provide; (3) a caveat line for the outcome panel stating measurability not causation; (4) pre-committed escalation thresholds; (5) a reading protocol with the practice-level-only rule; (6) a three-paragraph narrative template with the lane-specific grammar rules. Flag every threshold and cadence as a practice decision for me to set."
Then do the human work. Fill the row definitions verbatim from your Five-Metric Definition Sheet so the two documents cannot drift. Set thresholds from your own baseline and pilot data: Jordan's 80 percent latency threshold may be wrong for your payer mix. Confirm with your billing manager that the clearinghouse export actually distinguishes the reason codes the quality panel needs; many practices discover at month one that "denied" arrives as an undifferentiated lump, and fixing the export is a prerequisite, not a footnote. Decide the blended versus adopters-only display rule and write it down. And put the PHI rule at the top of page one: aggregates only, no client-level data, no note text, ever, so the page can travel to a board without a release.
Run the verification pass as a dry run: assemble one real month, last month, from the spec alone, the three owners pulling their own panels and one person assembling. Time it; if assembly takes more than half a day of combined effort, simplify the spec before institutionalizing the pain. Check that every number traces to a named report, the outcome caveat line printed, no client-identifying data appears anywhere, and the three narrative paragraphs obey the grammar. Then table-read it in fifteen minutes with the three-questions-per-panel protocol and note where the conversation tried to jump lanes.
Done looks like this: a spec any competent successor could execute, one assembled month attached as the worked exhibit, owners and deadlines named, thresholds pre-committed, the caveat line in place, signed by the clinical director and filed beside the definition sheet and the AI policy. Next lesson, this dashboard becomes the evidence base for the one-page ROI memo the board asked for.
Key Takeaways
- The monthly dashboard pulls four sources, EHR timestamps, AI vendor usage data, clearinghouse claims, and MBC platform data, onto one de-identified, practice-level page. The usage column distinguishes "the tool does not work" from "half the practice is not using it," two diagnoses with opposite remedies.
- Every number belongs to exactly one of three lanes: productivity (time and throughput, moves in weeks), quality (denials, documentation standard, MBC completion, moves in quarters), and outcomes (aggregated PHQ-9/GAD-7/PCL-5 trajectories, moves in quarters to years). Reading one lane as another is a category error a sophisticated board catches in seconds.
- The productivity panel pairs documentation minutes with signature latency and after-hours charting so speed cannot be gamed, and computes hours returned (Jordan's worked month: 10 minutes saved across 110 weekly sessions is about 18 hours a week) without converting to dollars; that conversion is the ROI memo's job, with stated assumptions.
- The quality panel separates AI-addressable denials (medical necessity, documentation insufficiency) from untouchable ones (eligibility, COB) and uses the untouchable series as a built-in control. The monthly ten-note audit catches what speed metrics cannot: fast, signed, thin notes that fail a 90837 review.
- The outcome panel prints its caveat on the panel itself: measurability and safety monitoring, never causation. Its operational rule is the tripwire: two consecutive quarters of cohort deterioration pauses expansion and triggers a clinical review. AI never scores a CSSRS or assigns a risk level anywhere on the dashboard.
- Pre-committed escalation thresholds (latency below 80 percent, sustained addressable-denial uptick, usage below 70 percent at month six, the two-quarter outcome rule) prevent both panic over noise and rationalization of signal. Numbers are read at practice level only; coaching belongs in supervision, never on the leadership screen.
- The three-paragraph monthly narrative enforces the taxonomy by grammar: time units for productivity, rates with controls for quality, monitoring language for outcomes. Twelve months of these paragraphs make the ROI memo a compilation; twelve months of bare numbers make it forensic reconstruction.
Skill.re