โ†
AI Agent Builders & Citizen Developers
Visionary ยท M8 ยท lesson 8 of 24 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
Hub-and-Spoke vs. Embedded vs. Hybrid: Picking the Operating Model
๐Ÿ“–
now learning

Hub-and-Spoke vs. Embedded vs. Hybrid: Picking the Operating Model

15 min

In Q1 2026, a global industrial company with 18,000 employees ran a six-month bake-off between two operating models for their agent program. The eastern hemisphere, led by a head of engineering excellence in Singapore, ran a pure hub-and-spoke model โ€” every agent designed by a central team in Bangalore, reviewed by a central governance board in London, shipped to business units who operated but did not modify. The western hemisphere, led by a head of digital in Houston, ran a fully embedded model โ€” every business unit hired their own builders, picked their own stack, owned their own agents end-to-end, with no central technical team at all. At the six-month review, the eastern hemisphere had shipped 4 production agents to a uniformly high standard and had a backlog of 31 more in queue. The western hemisphere had shipped 19 production agents at wildly varying quality, including two that had to be killed mid-quarter for compliance reasons. The CEO asked the chief AI officer for a recommendation. The recommendation was neither: it was the hybrid-with-shared-rails pattern, which by mid-2026 had become the consensus winner at scale. This lesson is the honest read of all three models โ€” what each actually delivers, where each actually breaks, and how to map the choice to the company's culture, regulatory exposure, and platform-team maturity.

Why This Choice Is Not About Elegance

The operating-model question gets attacked from three directions that are all wrong. The first is the elegance argument โ€” "embedded is more elegant because it pushes ownership to the edge." Elegance is not a property an enterprise agent program optimizes for; the same elegance argument would have made every Fortune 500 IT department a holacracy by 2018, and the ones that tried mostly reverted. The second is the consultant argument โ€” "best practice is to start with a center of excellence." Best practice in 2026 is whatever has not yet collapsed under its own contradictions in the past eighteen months, and pure CoE has collapsed in more programs than it has survived. The third is the technology argument โ€” "the right model depends on the agent framework we choose." Frameworks are easier to swap than org structures; the org structure is the longer-lived commitment.

The honest framing is that all three models โ€” hub-and-spoke, fully embedded, hybrid-with-shared-rails โ€” solve a different trade-off and impose a different cost. The strategist's job is not to pick the best one in the abstract. The strategist's job is to read the company's culture (top-down or bottom-up), the platform team's maturity (does one exist yet, and is it credible), the regulatory exposure (single-point-of-accountability requirement or not), and the political reality (who just announced what, who was just burned by which failure), and pick the model whose costs the company can actually bear at its current stage.

The most expensive operating model in 2026 is not any of the three. It is the unspoken hybrid where the company has implicitly chosen one model in some places and a different model in others, never reconciled them, and lets the contradiction generate quarterly political crises. The strategist who picks a model deliberately and writes it down โ€” even an imperfect one โ€” outperforms the strategist who waits for the perfect model and lets the contradictions accumulate.

Hub-and-Spoke: The Fast-Standards, Slow-Velocity Model

Hub-and-spoke is the operating model where a central hub team designs, reviews, and standardizes agents while business-unit spokes execute under the hub's direction. The hub typically sits under the CIO, CTO, chief AI officer, or chief digital officer. The hub has its own engineers, designers, PMs, and ops staff. The spokes are the business units; they receive the agents the hub builds, they operate them with hub support, and they can request changes through the hub's intake process. Variants of this pattern include the federal agency model (one strong center, several mostly-passive spokes), the consulting-firm model (the hub bills the spokes for delivery), and the platform-with-implementation-services model (the hub owns the platform and also does most of the building).

What hub-and-spoke is good at

The hub-and-spoke model adopts standards faster than any alternative. When the central team makes a decision โ€” to upgrade to Claude 3.5 Opus, to require eval coverage above 80%, to standardize on LangGraph for any multi-step orchestration, to deprecate a deprecated MCP server โ€” that decision is implemented across the agent portfolio in the time it takes the hub to do the work. There is no negotiation with twelve business units. There is no waiting for the slowest unit to comply. There is no version skew where the sales agent runs on this month's standards and the support agent runs on last quarter's. The hub decides, the hub implements, the standards adoption happens.

For programs with strong regulatory pressure (large-bank consumer lending, parts of healthcare and insurance, government work, defense), this property is not a nicety. It is a requirement. A regulator who issues a finding on a Tuesday and expects compliance evidence by Friday is satisfied by hub-and-spoke in a way that federated programs cannot match without sustained, expensive cross-team coordination.

What hub-and-spoke is bad at

Hub-and-spoke is slow at the only thing that ultimately matters: getting agents to where the work happens. Every new agent goes through the hub's intake queue. Every change goes through the hub's prioritization. Every disagreement between the hub's view of "good" and the spoke's view of "useful" gets resolved with the hub winning on technical merit and the spoke losing on local context. The result, observed across at least nine large hub-and-spoke programs that ran from 2023 to 2025, is the same pattern every time:

  • The hub's intake queue grows past nine months within the first year. Spokes that need an agent quickly start building their own outside the hub's view โ€” shadow agents in Zapier, Make.com, n8n, or vendor copilots that do not show up in the agent registry.
  • The hub's engineers, technically excellent, do not know the work deeply enough. The agents work in demo and underperform in production. The spoke is unhappy. The hub is unhappy. Both blame each other.
  • The hub becomes the single accountable party when an agent fails publicly. The hub absorbs the political damage. Other units stay insulated. The hub leader is eventually reorged or replaced.
  • The hub locks into a single orchestration framework, a single model provider, a single platform vendor. When the market moves (Anthropic ships a cheaper model, Google ships a faster model, an open-source framework leapfrogs the commercial one), the hub is too invested to switch. Spokes end up running on second-best technology because the hub already standardized.

By the eighteen-month mark, the hub-and-spoke programs in the cohort had standards adoption rates that any auditor would envy and production agent counts well below the embedded peers. The bake-off the industrial company ran in Q1 2026 โ€” 4 agents in the hub-and-spoke hemisphere, 19 in the embedded one โ€” is the typical pattern.

Fully Embedded: The High-Velocity, High-Drift Model

The fully embedded model puts the builders inside the business units with no central technical team at all. The sales unit hires its own agent builders. The support unit hires its own. The finance unit hires its own. Each unit picks its own tools, its own orchestrator, its own eval framework, its own observability stack, its own model providers. There is no shared platform. There may be a shared policy document and a shared governance forum, but the technical substrate is whatever each unit chose.

What fully embedded is good at

Fully embedded ships agents faster than any alternative. The builder sits next to the domain expert. The roadmap is the unit's roadmap. The eval set is the unit's actual workflow. The agent ships when the unit decides it ships, not when a central team's queue allows. The number of meetings required to ship an agent โ€” the metric that more than any other predicts how many agents a program ships in a year โ€” is dramatically lower in embedded programs than in hub-and-spoke ones. At the industrial company's western hemisphere, the median time from "we want an agent for this" to "the agent is in production" was 42 days. At the eastern hemisphere, it was 187 days.

Embedded also produces agents that perform better in their own context. The sales agent built by sales builders against sales eval data performs the sales job better than the sales agent built by a central team. The domain knowledge is encoded in the eval set, the prompt, the tool choices, the escalation paths. The fit between the agent and the work is higher because the people who do the work are closer to the building.

What fully embedded is bad at

Fully embedded drifts. By the end of year one, the company has five different eval frameworks (one team picked Braintrust, one picked LangSmith, one rolled their own, one is using spreadsheets, one has no eval set at all). It has three different observability stacks. It has six different orchestrators and twelve different model-call paths. When the EU AI Act's Article 26 deployer obligations bind in August 2026, the company has no consolidated way to demonstrate logging, oversight, or human-in-the-loop controls across the portfolio. When a customer DPA audit asks "which agents touch our data, where are they hosted, and what is the access pattern," the answer takes six weeks to assemble and is incomplete.

The drift problem is not solved by writing more policy documents. Policy documents do not enforce themselves. The drift is solved by shared infrastructure that makes the right path the easy path โ€” which is exactly what fully embedded programs do not have. The western hemisphere's two killed agents in the industrial company's bake-off were both compliance kills: one had been routing EU customer data to a US-region model endpoint because the unit's builder did not know about region pinning, and one had been auto-approving expense reports above the unit's policy threshold because the eval set did not include adversarial cases.

Fully embedded also makes the platform-economics broken. Every unit pays separately for its model gateway, its observability platform, its eval tooling. The same dollar spend buys substantially less capability than it would under a shared substrate. The CFO who runs the annual review notices the duplication and either forces consolidation (which the embedded units resist) or accepts a 30-50% cost premium for the privilege of not having a central team.

Hybrid-With-Shared-Rails: The 2026 Consensus Winner at Scale

The hybrid-with-shared-rails pattern is what the 2026 winners โ€” across financial services, technology, healthcare, industrials, retail, professional services โ€” converged on after attempting one or both of the alternatives. It is not the average of hub-and-spoke and fully embedded. It is a specific architectural opinion: the platform owns a defined set of rails (eval, observability, guardrails, identity, MCP allowlist, model gateway, registry, cost controls), and the business units own everything else (the agents, the prompts, the orchestration choice within supported options, the eval cases, the UX, the rollout, the on-call). Governance is overlay, not central building.

What the rails are and what they are not

The rails are the parts of the agent lifecycle where standardization compounds value and divergence creates risk. The non-rails are the parts where local context matters more than uniformity.

Rails โ€” owned by the central platform team:

  • Evaluation substrate (one harness: Braintrust, LangSmith, Inspect AI, or internal; one regression discipline; shared dataset patterns)
  • Observability stack (Langfuse, Helicone, Arize, Datadog LLM Observability; one trace UI; one sample-and-tag workflow)
  • Guardrails (Lakera, Protect AI, internal libraries; one PII redaction service; one prompt-injection screen)
  • Identity and authorization (Okta workload identities, SPIFFE, agent-scoped credentials, the policy-decision-point)
  • MCP allowlist (one catalog, one review process, one deprecation pipeline)
  • Model gateway (LiteLLM, Portkey, internal Bedrock/Azure OpenAI/Vertex gateway; one egress, one cost-attribution model, one kill switch)
  • Agent registry (one source of truth)
  • Cost controls (per-team budgets, per-call attribution, anomaly alerting)

Non-rails โ€” owned by the embedded builders inside business units:

  • The agents themselves (system prompts, tool selection within allowlist, orchestration choice within supported set)
  • The eval set (cases, golden answers, thresholds โ€” within the platform's harness)
  • The UX (chat surface, embedded widget, API contract)
  • The rollout plan (cohort, canary, full launch, sunset)
  • The on-call (the unit's own engineers respond to their own agent's incidents)
  • The roadmap (the unit's leadership prioritizes its own agent backlog)

The split is deliberate. Rails are where divergence creates risk that the company as a whole bears (a missing eval set is a regulator finding; an unreviewed MCP server is a security incident; a model-call path that bypasses the gateway is a cost overrun nobody can attribute). Non-rails are where uniformity creates risk that the local team bears (a central team writing prompts for the sales agent will write worse prompts than the sales builder; a central team owning the roadmap will deprioritize the unit's most-needed features).

Why hybrid-with-shared-rails won in 2026

The pattern won because it produces the velocity of embedded with the standards adoption of hub-and-spoke, without producing either the drift of pure embedded or the bottleneck of pure hub. When the platform team upgrades Claude 3.5 Opus, every agent gets the upgrade through the shared gateway without each unit doing the work. When a new MCP server is added to the allowlist, every builder gets access immediately. When governance issues a tier-2 review requirement, the platform automates 80% of the evidence collection and the builders supply the remaining 20%.

The pattern also handles the executive question โ€” "who is accountable when an agent fails?" โ€” cleanly. The unit that owns the agent is accountable for the agent. The platform is accountable for the rails. Governance is accountable for the review. Three named owners, no diffusion. In the hub-and-spoke model, accountability concentrates at the hub and the hub leader takes the political damage for failures across the portfolio. In the fully embedded model, accountability diffuses and nobody owns the cross-cutting failures. Hybrid-with-shared-rails gives the cleanest accountability map.

Where hybrid-with-shared-rails breaks

The pattern is not free. Two failure modes deserve naming.

Failure mode 1 โ€” the platform team is not credible. If the central platform team does not actually deliver paved roads that builders want to use, the rails become walls. Builders route around. Shadow agents accumulate. The model collapses into the worst of both worlds โ€” a hub that does not deliver, and an embedded layer that does not coordinate. The fix is hiring: the platform team must be staffed with engineers who have built agents in anger, not engineers who have read about them. The first six hires into the platform team are the most important hires the entire program will make.

Failure mode 2 โ€” the rails grow too thick. The platform team, after a year of building rails, starts believing more rails are better. They add a rail for prompt versioning. They add a rail for required UX patterns. They add a rail for mandatory documentation format. By month eighteen, the rails have become the hub-and-spoke model in disguise. The fix is a written charter: the platform team commits in writing to the specific set of rails, with a high bar (formal architecture review, sponsoring executive sign-off) for adding new ones. Rails are subtractive by default.

Matching the Model to the Company

The strategist's actual decision is not which model is best in the abstract. It is which model the specific company can absorb and operate. Five dimensions matter.

Dimension 1 โ€” Culture (top-down vs bottom-up)

Top-down cultures (centralized decision-making, strong corporate functions, history of centralizing innovation projects) absorb hub-and-spoke easily and resist pure embedded. They can also absorb hybrid-with-shared-rails if the platform team is positioned with executive air cover and the rails are framed as standards rather than as bottlenecks. Bottom-up cultures (autonomous business units, weak corporate functions, history of letting units pick their own tools) absorb pure embedded easily and resist hub-and-spoke. They can absorb hybrid-with-shared-rails if the platform is framed as a service to the units rather than as a control over them.

Dimension 2 โ€” Platform team maturity

If the company already has a credible platform-engineering organization (with a track record of shipping internal developer platforms, golden paths, paved roads for things like CI/CD or observability), hybrid-with-shared-rails works on day one because the platform team has the muscle. If the company does not โ€” if "platform team" means a name on the org chart with no track record โ€” the hybrid model requires a year of platform-team building before the rails are real. In the interim, a hub-and-spoke with explicit transition timeline is sometimes the right choice: stand up the hub, use it to build the rails over twelve months, then transition to hybrid as the rails mature.

Dimension 3 โ€” Regulatory exposure

Heavily regulated industries with single-point-of-accountability requirements (large-bank consumer lending under CFPB scrutiny, parts of healthcare under HIPAA and state-level health AI laws, regulated insurance lines, defense contracting with explicit single-accountable-party requirements) often need the formal accountability structure that hub-and-spoke provides. The practical solution is hybrid-with-shared-rails where the central function is the accountable party of record while the building work is federated. The org chart looks hub-and-spoke for regulatory purposes; the building model is hybrid for delivery purposes. Lightly regulated industries can run hybrid-with-shared-rails outright.

Dimension 4 โ€” Existing operating-model pattern

Companies that have decentralized software engineering over the past decade โ€” that have product teams owning their own services, their own on-call, their own roadmaps โ€” will absorb hybrid-with-shared-rails natively because the agent program looks structurally like the engineering organization. Companies whose software engineering is still centralized will find hub-and-spoke familiar and hybrid-with-shared-rails strange; the agent program may become the wedge that finally federates the engineering organization, but the strategist needs to know that is the political project they are signing up for.

Dimension 5 โ€” Political reality

If the CIO has just announced the launch of a new AI Center of Excellence with herself as executive sponsor, recommending fully embedded requires careful positioning and will probably fail; recommend hub-and-spoke with an explicit transition plan to hybrid-with-shared-rails over twelve to eighteen months. If the CIO was just burned by a centralization failure and is publicly looking for a federated approach, hybrid-with-shared-rails is the easy sell. If the chief AI officer is new and looking to establish a foothold, hybrid-with-shared-rails gives them a defensible central function (the platform team and governance) without making them the bottleneck for every agent.

The Diagnostic and the Decision Memo

The strategist who has to recommend a model needs a one-page diagnostic and a two-page decision memo. The diagnostic is the input; the memo is the output.

The diagnostic โ€” eight questions

  1. What is our headcount? Below 500: the choice is mostly cosmetic. 500-2,500: hybrid-with-shared-rails as target, hub-and-spoke acceptable as a transitional state. Above 2,500: hybrid-with-shared-rails is the only model that scales without producing the documented failure patterns.
  2. How regulated is our most exposed business? Heavily regulated (banking consumer lending, healthcare clinical, defense): formal-accountability variant of hybrid. Lightly regulated: hybrid outright.
  3. Do we have a credible platform-engineering organization today? Yes: hybrid works day one. No: hub-and-spoke with twelve-to-eighteen-month transition plan.
  4. Is our software-engineering organization centralized or federated? Federated: hybrid feels native. Centralized: hub-and-spoke feels native and hybrid will feel strange; budget political capital for the change.
  5. What is our cultural default โ€” top-down or bottom-up? Top-down: framing the platform as standards is important. Bottom-up: framing the platform as service is important.
  6. How many production agents do we have today, and how many do we expect in eighteen months? Under 5 today and under 20 expected: any model works. 20+ today or 50+ expected: hybrid-with-shared-rails or accept the documented failure modes.
  7. What did the CIO/CTO/chief AI officer just announce? A new CoE: recommend hub-and-spoke with explicit transition plan. A federated push: recommend hybrid-with-shared-rails as the disciplined version. Nothing yet: the strategist's recommendation will set the pattern.
  8. Who pays for the platform team in the first year? Central corporate budget: hub-and-spoke or hybrid both work. BU chargeback only: hybrid requires careful sequencing so the platform team is staffed before chargeback can fund it.

The decision memo โ€” five sections

  1. The recommendation. One sentence: "I recommend the hybrid-with-shared-rails operating model, with a twelve-month transition from our current state."
  2. The rationale. Three to five paragraphs tying the recommendation to the diagnostic answers. Name the specific company-state inputs that drove the choice.
  3. The alternatives considered. The other two models, with the specific reasons each was rejected for this company at this stage. Name the failure modes you are avoiding by not choosing them.
  4. The transition plan. If the recommendation requires moving from the current state, the rough quarterly milestones for getting there. Headcount changes, reporting-line changes, charter changes, communications plan.
  5. The trigger to revisit. What would cause the strategist to recommend a different model in twelve months? (Examples: platform team fails to deliver credible rails; regulatory environment changes; major reorg shifts the political baseline; an acquisition imports a different model that must be integrated.) The memo that names its own revision triggers is the memo that survives political attack.

Three Stories From 2025 and Early 2026

Story 1 โ€” The bank that picked hub-and-spoke and stayed there

A top-twenty global bank in early 2024 announced an AI Center of Excellence reporting to the chief data officer. By late 2025, the CoE had shipped 7 production agents (all internal-facing) and had a backlog of 84 documented requests from business units. The bank's regulatory exposure (a recent consent order on automated decisioning) made the formal-accountability structure non-negotiable. The CoE leader, when asked at a 2026 conference whether they would shift to hybrid, said the answer was yes but the transition would take three years because the platform-team capability did not yet exist and the regulator would not accept a transition until it did. The bank is running the right model for its current state, but the strategist knows the limitations and the transition plan.

Story 2 โ€” The SaaS company that picked fully embedded and pivoted at month fourteen

A 4,000-person SaaS company in mid-2024 went fully embedded. Each product team hired its own agent builders. By month twelve, the company had 23 production agents, four different eval frameworks, three observability stacks, and an annual model spend that nobody could attribute by team because each team had its own gateway. A customer DPA audit in month fourteen took six weeks to answer and surfaced two compliance issues that triggered a $1.2M settlement. The chief technology officer mandated a pivot to hybrid-with-shared-rails. The platform team was stood up over the following two quarters; by month twenty-four, the program looked structurally similar to the hybrid programs that had skipped the embedded phase, but it had paid roughly $4M in remediation and lost-velocity costs to learn the lesson.

Story 3 โ€” The industrial company that picked hybrid from day one

A 22,000-person industrial company in early 2025 hired a chief AI officer whose previous role had been running platform engineering at a tech company. She knew the hybrid pattern from outside the agent context. Her first ninety days were spent standing up the platform team (eight engineers, one PM, one SRE) and writing the rails charter (a four-page document specifying exactly which rails the team owned). The business units started hiring builders in month four. By month twelve, the company had 14 production agents across nine business units, a unified eval framework, and a governance council reviewing the portfolio quarterly. The chief AI officer's annual review noted that the program had spent 20% less on tooling than the bake-off projections predicted because shared infrastructure compounded value. The CEO renewed her mandate for year two with an expanded budget.

Key Takeaways

  • Three models, three trade-offs. Hub-and-spoke buys fastest standards adoption at the cost of BU velocity. Fully embedded buys highest velocity at the cost of highest drift. Hybrid-with-shared-rails is the 2026 consensus winner at scale because it produces embedded velocity with hub-grade standards adoption.
  • Hub-and-spoke failure pattern. Nine-month intake queue, shadow agents in spokes, hub blamed for portfolio failures, single-vendor lock-in, reorged within eighteen months.
  • Fully embedded failure pattern. Five eval frameworks, three observability stacks, no consolidated governance view, six-week regulator response, compliance kills from drift.
  • Hybrid-with-shared-rails opinion. Platform owns a defined small set of rails (eval, observability, guardrails, identity, MCP allowlist, model gateway, registry, cost controls). Business units own everything else (the agents, the prompts, the eval cases, the UX, the rollout, the on-call). Rails are subtractive by default.
  • The two hybrid failure modes. Platform team that is not credible (rails become walls, builders route around) and rails that grow too thick (hybrid collapses back into hub-and-spoke).
  • Five matching dimensions. Culture, platform-team maturity, regulatory exposure, existing engineering-org pattern, political reality. Map all five before recommending.
  • The diagnostic is eight questions. Headcount, regulatory exposure, platform-team maturity, engineering-org pattern, cultural default, agent count trajectory, recent executive announcements, who funds the platform.
  • The memo names its own revision triggers. The memo that says when it should be revisited survives political attack; the memo that does not, does not.
  • Standards-adoption speed is a real benefit of hub-and-spoke. If regulatory exposure makes that property non-negotiable, the right answer is hybrid-with-shared-rails with the central function as accountable party of record โ€” not pure hub-and-spoke.