AI for Financial Advisors & Wealth Managers
Visionary · M15 · lesson 15 of 18 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Piloting, Scaling, and Killing — A Practice's AI Innovation Discipline
📖
now learning

Piloting, Scaling, and Killing — A Practice's AI Innovation Discipline

15 min

The advisor practice that ships two proprietary AI workflows in 36 months wins. The practice that ships twelve, of which eight underperform and never get retired, ships nothing — it accumulates tech debt, splits attention across underperforming surfaces, and produces a buyer-diligence story that reads as 'experiments without endpoints.' Innovation discipline in 2026 is three habits: a 90-day pilot framework with explicit kill criteria, a scale decision that demands documented performance before promotion, and a retirement discipline that pulls underperformers before they accrete. This lesson installs all three for the senior advisor or Chief AI Officer running a $500M-$5B RIA's proprietary workflow portfolio — the L5 Ch2 L1 architecture meets its operating cadence.

Why Discipline Matters More Than Ideas

The Schwab 2026 RIA Benchmarking Study, the Kitces AdvisorTech map March 2026, the McKinsey 2025 wave on AI in professional services, the BCG advisor productivity research, and the practitioner reports from the named aggregators (Carson Group, Mariner, Cresset, Hightower, Captrust) all converge on one finding: pilot success rate at advisor firms hovers around 30-40%. Three out of five proprietary workflows that look promising at pilot kickoff fail to clear the 90-day scale decision — usually because the workflow does not produce the documented productivity gain, does not survive principal review under FINRA Rule 2210, does not generate Marketing Rule 206(4)-1 substantiable client outputs, or simply lands in a niche too narrow to justify the maintenance liability.

The firm that does not have an explicit kill discipline keeps the failed pilots running quietly. Each failed pilot accumulates: a vendor contract that doesn't auto-cancel, a prompt sub-library that requires periodic review under Rule 4511, a RAG vault content area that someone must maintain, a principal-review queue exception slot that someone must staff. At 12 failed pilots running quietly, the firm's AI tech debt approaches the budget allocated to the original 36-month transformation. The maintenance liability exceeds the value produced.

The Mercer Capital Q4 2025 and ECHELON Q3-Q4 2025 RIA M&A diligence framework reads exactly this signal. The L4 Ch8 L1 premium-attribute checklist scores 'documented workflows' and 'adoption metrics' as tier-determining; a firm with 12 half-deployed proprietary workflows scores worse than a firm with 3 fully-institutionalized workflows, even if the half-deployed count looks more impressive on a slide. Innovation discipline is the difference between the premium-tier 8x-10x multiple and the broader-market 6x-8x median.

The 90-Day Pilot Framework

A proprietary AI workflow pilot has a fixed 90-day duration. Not 60, not 120, not "we'll see how it goes." The 90-day duration is calibrated to four windows: enough time to deploy and stabilize (Days 1-30), enough time to capture real client-engagement use (Days 31-60), enough time to measure outcomes against the success criteria and produce the scale decision (Days 61-90).

Pilot Kickoff Deliverables (Day 0)

Before Day 1, the pilot has produced six explicit deliverables. (1) Hypothesis statement: what specific problem the workflow solves, for which client population, with what measurable outcome. ("This workflow reduces senior advisor time on a business-owner-exit engagement from 24 hours to 8 hours per engagement, while maintaining or improving the Reg BI documentation quality and producing client outputs that pass Marketing Rule 206(4)-1 principal review at first attempt 80% of the time.") (2) Pilot population: which advisors are running the pilot, which client engagements are in scope, what is excluded. (3) Success criteria: 3-5 measurable metrics with target levels and minimum-acceptable levels. (4) Kill criteria: 2-3 metrics or events that automatically end the pilot before Day 90 (e.g., principal-review queue rejection rate >25%, single Reg S-P 17 CFR Part 248 incident, advisor satisfaction below threshold for two consecutive weeks). (5) Scale decision criteria: what specifically must be true to promote to broader rollout. (6) Pilot sponsor + niche owner + AI Compliance Specialist signed off, with an explicit budget allocation, an AI Governance Committee ratification, and Smarsh / Global Relay archive coverage configured under FINRA Rule 4511 + SEC Rule 204-2.

Pilot Execution (Days 1-90)

Days 1-30 are deployment and stabilization. The workflow is provisioned on the three-layer architecture (enterprise LLM + RAG vault + prompt library) per L5 Ch2 L1. The pilot advisors are trained. The principal-review queue under Rule 2210 is configured for pilot-tagged outputs with elevated sampling (e.g., 100% review during pilot vs. risk-based sampling at scale). The ROI dashboard per L4 Ch5 L2 captures pilot-specific metrics. The AI Risk Register entry for the pilot is logged.

Days 31-60 are real-engagement use. The pilot advisors run the workflow on real client engagements (with client engagement-letter and ADV Part 2A AI disclosure in place per L1 Ch5.3). Outputs flow through the principal-review queue. Metrics accumulate. Advisor feedback is captured weekly. The kill criteria are monitored — any breach automatically triggers a 7-day cure window followed by termination if not cured.

Days 61-90 are measurement and decision. The pilot's metrics are tabulated against success criteria. The principal-review queue rejection rate is computed. The advisor-satisfaction score is captured. The Reg BI documentation quality is sampled by the AI Compliance Specialist. The Marketing Rule substantiation file is built per L4 Ch7 L1 for any client-facing claim the pilot produced. The AI Governance Committee receives the pilot review package at the Day-90 monthly meeting and votes on the scale-or-kill decision.

The Scale Decision — What Must Be True to Promote

A pilot that has produced positive metrics is not automatically scaled. The scale decision requires affirmative answers to seven questions, each of which the AI Governance Committee per L4 Ch6 L1 examines.

Question 1 — Did the pilot hit success criteria on at least 3 of the 4-5 named metrics? A pilot that produced productivity gain but failed on Reg BI documentation quality does not scale. A pilot that improved client satisfaction but failed on principal-review queue rejection rate does not scale.

Question 2 — Did the workflow produce zero unrecoverable failures? An unrecoverable failure is a Reg S-P 17 CFR Part 248 incident, a Marketing Rule 206(4)-1 violation that required substantive remediation, a Reg BI Care Obligation failure documented in a FINRA AWC-pattern memo, or an Investment Advisers Act of 1940 fiduciary duty breach (per the L1 Ch1 framing). One unrecoverable failure during pilot disqualifies scale.

Question 3 — Does the workflow integrate with the firm's existing supervisory architecture? The principal-review queue under Rule 2210 must absorb the workflow's outputs at scale-volume without becoming the bottleneck. The Rule 4511 retention pipeline must capture every artifact. The AI Risk Register must accommodate the workflow's risk profile.

Question 4 — Is the niche concentration sufficient to justify the maintenance liability at scale? A workflow that serves 12 households across the firm does not get scaled regardless of metrics; the maintenance liability exceeds the value. The threshold the lesson uses: at least 5% of firm households served by the workflow, or at least 15% of firm revenue attributable to the niche the workflow serves.

Question 5 — Is the senior-advisor niche-owner persistently engaged? A workflow whose niche owner is one quarter from retirement, or who is the sole bearer of the prompt library institutional knowledge, fails the talent/key-person diligence dimension per L4 Ch8 L1. The committee requires evidence of cross-training and prompt-library succession before scale.

Question 6 — Does the Marketing Rule 206(4)-1 substantiation file support the workflow's external claims? Any external claim the firm wants to make about the workflow ("our proprietary business-owner-exit workflow") must be substantiable per L4 Ch7 L1 and consistent with the January 2026 SEC staff FAQs framing + the 2024-2025 Delphia/Global Predictions AI-washing precedent.

Question 7 — Does the scale-cost projection still fit the budget envelope? Pilot costs are often 30-40% of scale costs because pilot uses elevated sampling, manual review, and dedicated principal time. Scale projections must update the L4 Ch5 L2 ROI dashboard with realistic scaled-volume costs.

Seven affirmative answers = scale. Six = scale with conditions documented in committee minutes. Five or fewer = kill or extend the pilot one cycle with explicit remediation plan.

The Kill Discipline — Retiring Underperforming Workflows

The harder of the three habits is killing underperforming workflows. Pilots that fail get killed easily; the harder kill is the workflow that scaled, ran for 18 months, and is now visibly underperforming. The firm's natural inertia preserves it: vendor contracts are paid, prompt library exists, advisors trained, a small population still uses it. Killing it requires a deliberate trigger.

The Quarterly Portfolio Review

The AI Governance Committee reviews the proprietary workflow portfolio quarterly — not monthly (too frequent for stable workflows) and not annually (too infrequent to catch decay). The quarterly review applies five health metrics to every active workflow: utilization (number of distinct client engagements in the quarter), advisor satisfaction (the named-niche-owner score plus the broader-advisor-user score), principal-review queue rejection rate (compared to baseline), maintenance burden (hours per quarter the Prompt Librarian + AI Compliance Specialist + niche owner spend on this workflow), and revenue/engagement-value contribution (workflow-attributable revenue or workflow-attributable engagement value).

A workflow that scores below threshold on 2 of 5 metrics for two consecutive quarterly reviews triggers the kill review. The kill review is a one-page memo by the AI Compliance Specialist documenting: (a) what the workflow does, (b) the four-quarter metric trend, (c) the recommendation (kill, narrow scope, reinvest, transfer to a different niche owner), (d) the migration plan if killed.

The Migration Plan When Killing

Killing a workflow is not just turning it off. The migration plan covers six items: (1) Active engagement transition: any open client engagements running on the workflow get completed under the existing workflow with archived artifacts under Rule 4511, or transitioned to an alternative workflow with documented Reg BI Care Obligation rationale for the change. (2) Vendor contract cancellation: any vendor contracts specific to the workflow are cancelled, with attention to auto-renewal windows. (3) Prompt library archival: the workflow's prompts move from active library to archived library under Rule 4511 retention — the prompts remain records even after retirement. (4) RAG vault content disposition: workflow-specific content stays in the vault but is tagged as deprecated; client NPI is purged per Reg S-P 17 CFR Part 248 retention schedules. (5) Advisor communication: the affected advisor pods are notified and trained on the alternative. (6) External communication review: any external Marketing Rule 206(4)-1 claim about the workflow is removed from the firm's website, ADV Part 2A, marketing materials, and the substantiation file is updated; the L4 Ch7 L1 audit log captures the change.

The Portfolio Management Mindset

The senior advisor or Chief AI Officer running a proprietary workflow portfolio in 2026 is operating like a portfolio manager — diversified across 2-4 workflows in defended niches, rebalancing quarterly, harvesting losses (killing underperformers), and reinvesting in the strongest performers. The portfolio is not a museum of every workflow the firm has ever tried; it is a working portfolio with active positions and explicit position-sizing.

The position-sizing question gets explicit attention. A workflow serving 40% of the firm's revenue niche (the Austin case study's business-owner exit) gets disproportionate maintenance investment — additional prompt library iteration, dedicated niche-owner attention, the principal-review queue prioritization. A workflow serving 8% of revenue might run with minimal maintenance, accepting slower iteration in exchange for lower overhead. The L5 strategist's portfolio thinking is what produces the buyer-diligence story that scores 9-10 on the L4 Ch8 L1 premium-attribute checklist.

The 2026 RIA M&A diligence framework explicitly rewards portfolio thinking. The Mercer Capital and ECHELON Q3-Q4 2025 data show that practices with documented portfolio discipline — clean prioritization, explicit kill records, evidence of rebalancing — price at the top of the 8x-10x premium tier or premium-top ~11.6x. Practices with sprawling, unmanaged, half-deployed workflows price at the bottom of the broader-market 6x-8x median. The L4 Ch8 L1 5-dimension diligence reads the discipline directly.

Case Study — Killing a Workflow That Ran for 14 Months

A $2.4B RIA based in Boston deployed a proprietary expat tax-and-estate workflow in Q4 2024. The pilot cleared the 90-day scale decision in Q1 2025 with strong metrics (6 of 7 scale criteria affirmative; the niche owner had documented succession). Through 2025, the workflow served 38 engagements and produced $580K of attributable revenue against $190K of all-in cost — a 3x payback in year one.

In Q3 2025, two shelf vendors entered the expat space with credible products. By Q1 2026, one of them (a partnership between a Big-4 international tax practice and a wealth-tech firm) was shipping a SaaS product at $4,800 per advisor seat per year that covered 80% of the firm's proprietary workflow scope. The firm's Q1 2026 portfolio review surfaced the question: should we kill our proprietary workflow and adopt the shelf alternative?

The kill memo covered the six migration items above. The committee voted to kill. Reasoning: (a) shelf-vendor 80% scope coverage at $4,800/seat × 14 advisor seats = $67K/year vs. $190K/year proprietary maintenance — direct cost savings; (b) the 20% gap was in country-pair treaty interpretation, which the firm's outside counsel was already delivering on a per-engagement basis; (c) the proprietary workflow's principal-review queue rejection rate had drifted from 12% in Q2 2025 to 19% in Q1 2026 (decay); (d) the niche-owner senior advisor was now devoting time to a new pilot in private-fund accredited-investor onboarding, a higher-leverage frontier.

The migration: 6 active engagements completed under the proprietary workflow with Rule 4511 archived; 12 future engagements onboarded to the shelf vendor with documented Reg BI Care Obligation rationale; vendor contracts for the proprietary infrastructure layer scaled down; prompt library archived under Rule 4511; ADV Part 2A AI disclosure updated to reflect the shelf-vendor relationship per Reg S-P 17 CFR Part 248 vendor oversight; Marketing Rule 206(4)-1 substantiation file updated to remove the "proprietary expat workflow" external claim. The kill was complete within 60 days. The freed budget redeployed into the new pilot.

The diligence outcome: in the firm's Q2 2026 shadow buyer diligence, the kill was scored as a positive signal under the L4 Ch8 L1 documented-workflows and adoption-metrics dimensions. Evidence of disciplined kill increased the firm's premium-attribute score from 8/10 to 9/10. The lesson: killing a workflow that ran for 14 months is not failure; it is the discipline that distinguishes a portfolio from a museum.

The Anti-Patterns to Avoid

Three anti-patterns recur in the practitioner reports and the post-mortem analyses on stalled AI transformations.

Anti-pattern 1 — The Pilot That Never Ends: a pilot that ran past 90 days into 180 days, then 270 days, then quietly became "production." No scale decision was made; the workflow accumulated maintenance without ever being held to the seven scale criteria. The fix: enforce the 90-day clock; pilots that aren't ready for scale at Day 90 get killed or restarted as a new pilot with revised criteria.

Anti-pattern 2 — The Workflow Too Beloved to Kill: a senior advisor's pet project that produced minor revenue but no one wanted to retire because the advisor cared about it. The fix: portfolio review on metrics, not sentiment; the kill memo is a one-pager that the AI Compliance Specialist authors and the AI Governance Committee votes on — not the niche owner's call alone.

Anti-pattern 3 — Killing Without Migration: workflow is killed, vendor contracts cancel, advisors stop using it — but active engagements are abandoned, prompt library disappears (Rule 4511 violation), ADV Part 2A still describes the workflow, Marketing Rule substantiation file still references the workflow. The fix: the six-item migration plan is mandatory; the kill is not complete until the migration is documented and the AI Compliance Specialist signs the closure memo.

Key Takeaways

  • Pilot success rate at advisor firms hovers at 30-40%, per the Schwab 2026 study, Kitces AdvisorTech map, and McKinsey/BCG practitioner research. Three of five pilots that look promising fail the scale decision — usually on principal-review queue rejection rate, Reg BI documentation quality, Marketing Rule 206(4)-1 substantiability, or niche concentration. Without explicit kill discipline, failed pilots accumulate as tech debt.
  • The 90-day pilot framework is fixed-duration with six Day-0 deliverables: hypothesis statement, pilot population, success criteria (3-5 metrics), kill criteria (2-3 metrics or events triggering early termination), scale decision criteria, and signed sponsorship from niche owner + AI Compliance Specialist + AI Governance Committee + Smarsh / Global Relay archive configuration under FINRA Rule 4511 + SEC Rule 204-2.
  • The scale decision requires seven affirmative answers: hit 3-of-4-5 success metrics, zero unrecoverable failures (Reg S-P, Marketing Rule, Reg BI Care, fiduciary), supervisory-architecture integration, niche concentration (≥5% households or ≥15% revenue), engaged niche owner with cross-training, Marketing Rule substantiation, scale-cost projection within budget. Seven = scale; six = scale with conditions; five-or-fewer = kill or restart.
  • The kill discipline runs quarterly on five health metrics — utilization, advisor satisfaction, principal-review rejection rate, maintenance burden, revenue/engagement-value contribution. Two-of-five below threshold for two consecutive quarters triggers the kill review; the kill memo is a one-pager from the AI Compliance Specialist; the migration plan covers active engagement transition, vendor cancellation, prompt library archival under Rule 4511, RAG vault disposition under Reg S-P, advisor communication, and external Marketing Rule 206(4)-1 communication review.
  • The portfolio management mindset: 2-4 active workflows in defended niches with explicit position-sizing — disproportionate maintenance to the workflow serving 40%+ of niche revenue, minimal maintenance to the workflow serving 8% of revenue. Quarterly rebalancing. The Mercer Capital Q4 2025 and ECHELON Q3-Q4 2025 buyer-diligence framework rewards portfolio discipline directly — clean kill records and rebalancing evidence score at top of the 8x-10x premium tier; sprawl scores at broader-market 6x-8x median.
  • The three anti-patterns: the pilot that never ends (no scale decision, quiet drift to production), the workflow too beloved to kill (sentiment over metrics), killing without migration (Rule 4511 violation, ADV currency failure, Marketing Rule substantiation drift). Each has a named fix: enforce the 90-day clock, make kill decisions on metrics, mandate the six-item migration plan.