Data Governance and Integration Strategy
The L5 Ch3 L1 enterprise AI policy answers the question "what rules govern AI use at our firm?" This lesson answers the harder pair of questions underneath: who owns the client data, and where does it live? The L4 Ch1-L4 Ch4 strategic stack assumed AI tools read client data from the firm's existing systems (Schwab, Fidelity, Pershing, BNY Mellon custodians; RightCapital, eMoney, MoneyGuidePro planning platforms; Wealthbox, Salesforce FSC, Redtail CRMs; Smarsh, Global Relay archives) without examining whether the data architecture itself supports AI at scale. By the time a 200-advisor aggregator is at Wave 4 of the L5 Ch1 L1 transformation, the data architecture is the constraint. This lesson walks through the data-mesh-versus-data-lake decision, the consent management framework, the where-does-the-data-live question across custodian / planning / CRM / archive, and the Reg S-P 17 CFR Part 248 / GLBA Safeguards implications of each architectural choice. The Chief Information Officer owns this lesson; the Chief AI Officer, CCO, and CTO each have specific dependencies on the answer.
Who Owns the Client Data โ A Question With Five Stakeholders
The casual answer "the client owns their data" is correct as a matter of right but incomplete as a matter of architecture. Five stakeholders have operating claims on the client data layer and the architecture must reconcile all five.
The client โ has ownership rights under Reg S-P 17 CFR Part 248, GLBA Safeguards, state CPRA / DIR / equivalent privacy regimes, and the firm's engagement letter. The client controls consent to use beyond the engagement scope. The custodian (Schwab, Fidelity, Pershing, BNY Mellon, TD Ameritrade legacy migrated to Schwab-Advisor) โ holds the system-of-record for accounts, positions, transactions, beneficiary designations, and tax basis. The custodian's data-feed terms govern what the firm can re-export, re-process, and re-share. The advisory firm โ has fiduciary duty under the Investment Advisers Act of 1940 to maintain accurate client records, ADV Part 2A delivery records, IPS documents, Reg BI files, meeting notes, prompts-as-records under FINRA Rule 4511 + SEC Rule 204-2. The firm is the operating data steward. The vendor (planning software, CRM, archive, AI tools) โ accesses client data under data-handling agreements; the May 2024 Reg S-P 17 CFR Part 248 amendments require explicit vendor oversight. The regulator โ has examination rights to records that touch the firm's regulated activities. The data architecture must produce records on demand for SEC, FINRA, state DOI, NY DFS examiners.
The five-stakeholder reconciliation is the data governance problem. An architecture that optimizes for any one stakeholder at the expense of another produces operational friction or regulatory exposure. The lesson's framework optimizes for the firm's fiduciary duty while preserving all five stakeholders' operating claims.
Where Does the Data Live โ The Four-System Map
Every advisor practice's client data lives across four systems by default. Understanding what lives where is prerequisite to any architectural decision.
Custodian โ The System of Record
Schwab, Fidelity, Pershing, BNY Mellon hold the canonical account data: account number, registration, position-level holdings, transactions, cost basis (post-1099-B), dividends, interest, capital gains distributions, beneficiary designations, ACATs status, NIGO history. Custodian feeds (typically daily) push this data to the firm's planning and CRM systems. The custodian's terms of service govern re-use โ most custodians prohibit selling data, prohibit sharing with non-affiliated parties without consent, and require firms to maintain Reg S-P-equivalent safeguards.
Planning Software โ The Modeling Layer
RightCapital, eMoney, MoneyGuidePro hold the modeled view: goals (retirement, education, legacy), Monte Carlo assumptions and outputs, projected cash flows, Social Security claiming optimization, tax projections, asset-liability matching, IPS allocation targets. The planning software ingests custodian feeds for current positions and the advisor's manual inputs for goals and assumptions. Holistiplan and FP Alpha sit alongside planning software with extracted tax-return and estate-document data feeding into the planning layer or staying in their own systems.
CRM โ The Relationship Layer
Salesforce FSC + Einstein, Wealthbox, Redtail (including Redtail Engage), Practifi, Pulse360 hold the relationship view: contacts, households, activities, opportunities, tasks, meeting notes, communications history, document attachments, custom fields. The CRM AI layer reads household state and produces next-best-action prompts, summary briefings, activity drafts. The CRM is the firm's day-to-day operating system.
Archive โ The Recordkeeping Layer
Smarsh, Global Relay hold the immutable retention layer: emails, text messages, social media, meeting recordings and transcripts (via Jump or Zocks integration), Slack and Teams messages, prompts-as-records, AI outputs, edits, signoff trails. The archive operates under FINRA Rule 4511 and SEC Rule 204-2 retention obligations. AI Risk Register entries, policy amendments, vendor approval records, training records, IRP playbooks โ the recordkeeping layer holds the full operational history.
Data-Mesh Versus Data-Lake โ The Architectural Decision
The 2024-2026 enterprise data architecture conversation produced two competing models. Each has trade-offs that map directly to the firm's AI strategy and the L5 Ch3 L1 enterprise policy.
The Data-Lake Architecture
The data-lake model: a single centralized storage layer holds all client data extracted from custodian, planning, CRM, and archive systems. AI tools query the lake. Advantages: single point of integration; uniform data access patterns; easier ROI dashboard instrumentation; cleaner buyer-diligence story per L4 Ch8 L2 (one architecture to diligence). Disadvantages: single point of failure under Reg S-P 17 CFR Part 248; concentrated NPI risk; consent management complexity if the lake serves multiple firms in an aggregator network; vendor oversight burden if the lake itself is vendor-hosted; ADV Part 2A material change every time the lake's data scope changes; GLBA Safeguards single point of attack surface.
The Data-Mesh Architecture
The data-mesh model: client data stays in source systems; AI tools query each system through controlled interfaces with consent enforcement at the interface. Advantages: distributed risk; consent enforced at point of access; aligns with custodian and vendor terms-of-service that prohibit re-export; per-system Reg S-P 17 CFR Part 248 vendor oversight scales naturally; ADV Part 2A material change only when interface changes, not when data scope changes; matches the L5 Ch1 L1 hybrid model's distributed pattern. Disadvantages: more complex integration layer; harder to instrument cross-system queries; ROI dashboard requires federation; buyer diligence requires explaining the mesh.
The Hybrid Pattern โ What 2026 Aggregators Actually Ship
The hybrid pattern dominates 2026 production architectures at $1B-$20B RIA networks. Operational data stays in source systems (custodian, planning, CRM) per the mesh model; aggregated firm-internal data (compliance-blessed IPS templates, Reg BI memo library, approved client-facing concept memos, WSPs, ADV Part 2A history, prior 12 months of Marketing Rule audit substantiation files, prior 24 months of approved client-facing communications) lives in the firm's RAG vault per L5 Ch2 L1 โ a partial data lake for firm-internal content but explicitly not for client NPI. Client NPI never enters the firm-wide vault; it stays in source systems and is fetched per-engagement under engagement-letter and ADV Part 2A consent.
Consent Management Across the Architecture
Consent management is the single most-underbuilt layer in 2026 advisor AI architectures. The Reg S-P 17 CFR Part 248 May 2024 amendments tightened the vendor-oversight obligations; the engagement letter and ADV Part 2A disclosure paragraph are the consent vehicles; but operational enforcement of consent at the data-access layer is where firms most often fall short.
Engagement-Letter Consent
The engagement letter contains the broad consent for AI use in support of the advisory relationship: meeting capture, summarization, document extraction, draft generation, pattern detection, and ancillary administrative use. The standard 2026 language: "Client consents to Firm's use of AI tools to support the advisory relationship, including meeting capture, document extraction, draft generation, pattern detection, and other workflow support. Firm's use of AI tools is governed by Firm's Privacy Policy, ADV Part 2A, and the AI Use Policy. Client may request additional information about specific tools at any time. Specific advisor recommendations remain the responsibility of the registered individual."
ADV Part 2A Disclosure
The ADV Part 2A AI disclosure paragraph per L5 Ch3 L1 names the AI tool categories, anchors recommendations to registered individuals, and references Reg S-P 17 CFR Part 248 client-information handling. The annual updating amendment per L5 Ch7 L6 incorporates AI tool changes; off-cycle prompt amendments under IA-1992 cover material changes mid-year.
Opt-In Versus Opt-Out by Capability
The L5 Ch3 L3 framework develops the 2027-2028 trajectory in detail, but the consent baseline for 2026 is: agentic AI (action-taking AI per L4 Ch3 L3) is opt-in by client; supporting AI (meeting capture, document extraction, draft generation) is the firm's operating model with engagement-letter and ADV Part 2A disclosure, and the client may request specific tool exclusion ("do not use meeting AI in my reviews") which is captured in the CRM custom field and operationalized at the data-access layer. The opt-out request is itself a Reg BI Care Obligation consideration documented in the file.
Operational Enforcement at the Data-Access Layer
The data-mesh interface (or the data-lake access controls in a lake architecture) enforces consent at every AI tool query. When the Chief AI Officer's prompt against the RAG vault would touch a household with an opt-out flag, the interface either blocks the query or excludes that household's data from the response. The opt-out flag lives in the CRM (Salesforce FSC custom field, Wealthbox custom field, Redtail custom field) and propagates through the integration layer. The AI Risk Register tracks opt-out density as a pattern signal.
Custodian Feed Implications โ The Multi-Custodian Aggregator's Reality
An aggregator running across three custodians (e.g., Schwab as primary, Fidelity for institutional, Pershing for broker-dealer affiliated) faces three custodian-data-feed contracts, three sets of re-use restrictions, and three reconciliation patterns. The data governance must handle: (a) which custodian feed updates which household's data in the planning and CRM systems; (b) how cross-custodian aggregation occurs for households with assets at multiple custodians (the 2026 reality for most $1M+ households); (c) what the data-feed terms allow the firm to do with extracted data (most custodians prohibit reselling, prohibit sharing with non-affiliates without consent, require Reg S-P-equivalent safeguards downstream).
Practical pattern: the aggregator's data architecture treats each custodian feed as a Reg S-P 17 CFR Part 248 vendor relationship under Section 2 of the enterprise policy; SOC 2 Type II reports reviewed annually; data-handling agreements include explicit re-use scope; the L4 Ch6 L1 AI Risk Register tracks custodian-feed-specific risks (feed reliability, schema changes, contract changes affecting AI-system access).
Reg S-P and GLBA Implications by Architecture
Data-Lake Reg S-P Profile
A centralized data lake holding NPI from a 200-advisor network's 30,000 households is a Reg S-P 17 CFR Part 248 attack surface concentrating risk; the May 2024 amendments' 30-day breach notification could affect all 30,000 households simultaneously if a single incident occurs; GLBA Safeguards Rule single-point-of-failure exposure; the L4 Ch4 cybersecurity playbook becomes the critical defense. The 2026 winning aggregators avoid this pattern for NPI.
Data-Mesh Reg S-P Profile
A mesh keeping NPI in source systems distributes the Reg S-P attack surface; an incident at one source system affects only that system's data; the 30-day notification window scope reduces; GLBA Safeguards risk distributes; each source-system vendor has its own attack surface and own Reg S-P 17 CFR Part 248 vendor oversight. The mesh pattern is preferred for NPI.
Hybrid Reg S-P Profile
The hybrid pattern (mesh for NPI + lake for firm-internal content) inherits the mesh's NPI risk distribution and adds a separate Reg S-P-relevant attack surface only at the firm-internal vault (which by design does not hold client NPI). This is the structurally cleanest pattern under Reg S-P 17 CFR Part 248 and is consistent with the L5 Ch1 L1 hybrid model's broader architectural framing.
Case Study โ Aggregator Data Architecture Rebuild
A $9B aggregator with 165 advisors across 14 member firms initiated a data architecture rebuild in Q2 2025. The starting state was a partial data lake holding 6 months of cached custodian feed data plus 18 months of meeting-AI transcripts and 9 months of planning-software exports. The CCO had been raising Reg S-P 17 CFR Part 248 concerns since the May 2024 amendments โ the lake held NPI for 22,000 households across the network with consent management at the engagement-letter level but no operational enforcement at the data-access layer.
The rebuild over Q3 2025-Q1 2026: data-mesh architecture for client NPI with consent enforcement at the integration layer; firm-internal RAG vault for compliance-blessed content (IPS templates, Reg BI memo library, ADV Part 2A history, Marketing Rule substantiation files) explicitly excluding client NPI; per-source-system Reg S-P 17 CFR Part 248 vendor oversight with SOC 2 Type II reports for each of Schwab, Fidelity, Pershing custodian feeds, RightCapital, eMoney, MoneyGuidePro planning feeds, Salesforce FSC, Wealthbox CRM feeds, Smarsh archive feed; opt-out flag propagation through CRM custom fields; AI Risk Register entries for each integration point.
The Q2 2026 outcomes: zero Reg S-P 30-day breach notifications across the network (the prior 18 months had three close-call incidents, all in the lake); 22,000 households now operationally consent-enforced at the data-access layer; the L4 Ch8 L1 supervisory-architecture score increased 8/10 โ 9/10 on the data-handling dimension; outside counsel review at policy amendment found the architecture aligns with the May 2024 Reg S-P amendments, GLBA Safeguards, NY DFS 23 NYCRR 500, and the NAIC AI Model Bulletin third-party-service-provider expectations. The L5 Ch7 L6 AI-diff workflow for ADV Part 2A captured the architecture change as a material amendment filed off-cycle (prompt) under IA-1992.
Key Takeaways
- Five stakeholders own claims on the client data layer: the client, the custodian (Schwab, Fidelity, Pershing, BNY Mellon), the advisory firm, the vendor (planning software, CRM, archive, AI tools), the regulator. Data governance reconciles all five โ none can be optimized away.
- Four systems hold client data: custodian (system of record), planning software (modeling layer), CRM (relationship layer), archive (recordkeeping layer per FINRA Rule 4511 + SEC Rule 204-2). Each system has its own re-use restrictions, vendor terms, and Reg S-P 17 CFR Part 248 vendor oversight.
- The architectural decision is data-lake vs data-mesh vs hybrid. Pure lake concentrates Reg S-P attack surface โ 200-advisor network's 30,000 households one breach away from network-wide 30-day notification. Pure mesh distributes risk but adds integration complexity. The 2026 winning hybrid: mesh for client NPI (per source system) + firm-internal RAG vault for compliance-blessed content (explicitly no NPI).
- Consent management has three layers: engagement-letter (broad AI use consent), ADV Part 2A AI disclosure (regulatory disclosure of tool categories per L5 Ch3 L1), and operational enforcement at the data-access layer (opt-out flag in CRM custom field propagating through integration). Agentic AI is opt-in per L5 Ch3 L3 framework; supporting AI is operating model with engagement-letter + ADV disclosure plus client-requested tool exclusion.
- Multi-custodian aggregators handle each custodian feed as a Reg S-P 17 CFR Part 248 vendor relationship under Section 2 of the enterprise policy; SOC 2 Type II reports annually; explicit re-use scope in data-handling agreements; cross-custodian aggregation operationalized at the integration layer with consent enforcement.
- Case study: $9B aggregator with 165 advisors rebuilt from partial data lake to hybrid architecture across Q3 2025-Q1 2026. Outcomes: zero Reg S-P 30-day breach notifications post-rebuild (three close-call incidents pre-rebuild), 22,000 households operationally consent-enforced, L4 Ch8 L1 data-handling supervisory score 8/10 โ 9/10, architecture aligns with May 2024 Reg S-P amendments + GLBA + NY DFS 23 NYCRR 500 + NAIC AI Model Bulletin, ADV Part 2A material amendment filed off-cycle under IA-1992 per L5 Ch7 L6 AI-diff workflow.
Skill.re