Redesigning It Teams For Ai
Hook
Your help desk has a 4-hour average wait time for password resets. Your NOC spends two hours every morning manually reviewing logs to spot anomalies. Your infrastructure team is so busy managing servers that they haven't had time to upgrade to the latest OS. And you just got the message that three of your best network engineers have left to join competitors who are already AI-native.
Here's what they saw that you might not: your organization's shape is optimized for a world that no longer exists.
For the past 20 years, IT organizations have been shaped by constraints: the cost of storage meant centralizing data. The complexity of infrastructure meant specialization (DBAs, network engineers, systems engineers in separate silos). The friction of change meant stable teams with predictable responsibilities. AI changes all of this.
AI doesn't just change what IT does. It changes how IT is organized. Tasks that required humans can now be automated or augmented. Teams that were separate can now integrate. Specialized knowledge can now be democratized. But this only happens if you redesign the organization intentionally. If you keep the old structure and just add AI on top, you'll get neither the benefits of AI nor the benefits of the traditional structure. You'll get confusion, duplication, and cost.
This lesson is about the organizational redesign required to actually capture AI's benefits. It's not about adding headcount. It's about reshaping headcount for a different era.
Purpose
You will understand which IT teams grow in the AI era, which transform, which disappear, and which merge. You will learn the organizational design patterns that actually work, and you will have a framework for managing the transition without breaking current operations.
Why This Matters
Your organizational structure is a constraint on your strategy. If your help desk is structured as a cost center where specialists answer reactive tickets, you will never build an AI-augmented support system. If your NOC is organized around shift work and manual monitoring, you will never build autonomous operations. If your data team is isolated from your infrastructure team, you will never build the integrated data platform that AI systems require.
More specifically: the talent market is choosing. Right now, your best engineers are either moving to companies that are reorganizing for AI, or they're getting frustrated and leaving the industry. If you don't reorganize, you'll lose them. If you do reorganize well, you'll attract the engineers who want to work at the frontier of what IT can be.
And financially: misaligned organizations waste money. If help desk, NOC, and application support are all separate teams using different tools, you're paying for AI solutions three times instead of once. If you reorganize around capability, you invest in AI infrastructure once and leverage it everywhere.
Core Concepts
Key Insight: Some Teams Grow, Some Transform, Some Merge
The future IT organization will have more people in some roles and fewer in others. The shape will be radically different.
Teams that grow:
- AI Platform Engineering (new role)
- Data Engineering (grows 3-5x)
- AI Security and Governance (new role, grows)
- Automation Engineering (transforms from DevOps into something broader)
Teams that transform:
- Help Desk → AI-Augmented Support (same people, different tools, different skill mix)
- NOC → Autonomous Operations Platform (not gone, but run by AI with human oversight)
- Database Administration → Data Platform Engineering (DBAs are now platform builders, not firefighters)
- Network Operations → Autonomous Network Management (fewer people managing more with AI agents)
Teams that merge:
- Data Engineering + Infrastructure Engineering → Data Infrastructure (data pipelines and systems infrastructure converge)
- Security + Governance + Audit → Risk and Compliance (traditional silos don't work in AI era)
- Systems Administration + DevOps → Infrastructure Automation (sysadmins become cloud engineers, DevOps becomes infrastructure platforms)
Key Insight: The Hub-and-Spoke Model vs. Fully Embedded vs. Centralized Platform Model
There are three primary organizational design patterns for AI in IT operations:
Model 1: Fully Centralized Platform Team
- One central AI Platform team owns all AI infrastructure
- All AI applications go through this team
- Pros: consistency, efficiency, strength
- Cons: slow, bottleneck-prone, lacks domain knowledge
- Works for: organizations with 5-20 active AI workloads
Model 2: Hub-and-Spoke
- Central platform team provides "spokes" (shared infrastructure)
- Individual teams (Help Desk, NOC, Infrastructure) have embedded engineers who use the platform
- Pros: both speed and consistency, domain knowledge at the edges, but leverage at the center
- Cons: more expensive, requires more coordination
- Works for: medium to large organizations with 20-50 active AI workloads
Model 3: Fully Distributed with Light Governance
- Every team builds their own AI solutions
- Centralized governance and security team sets standards
- Pros: maximum speed, teams move independently
- Cons: massive duplication, drift, security risk
- Works for: technology companies with exceptionally strong engineers and risk tolerance
- Warning: Usually reverts to Model 2 after the first incident
Recommendation: Start with Model 1 (centralized platform), move to Model 2 (hub-and-spoke) as you scale.
Key Insight: Help Desk 2.0 Is Not Help Desk 1.0 with a Chatbot
Most organizations' first AI initiative in IT operations is a help desk chatbot. They bolt a chatbot onto the existing help desk structure and call it done. What actually happens: the chatbot answers 10% of tickets, help desk staff don't know how to work with it, and nothing changes operationally.
The real transformation requires rethinking help desk work:
- Level 1 (Chatbot/AI) answers password resets, account unlocks, software requests, and knowledge base questions. Human review happens in the background.
- Level 2 (AI-Augmented Humans) handles more complex issues. The human has AI summarizing ticket history, suggesting solutions, and documenting resolutions.
- Level 3 (Escalation) handles the truly complex issues.
But this only works if you:
- Retrain help desk staff to be AI operators, not ticket answerers
- Change their KPIs from tickets answered per hour to customer satisfaction and issue resolution
- Build the AI infrastructure and knowledge base first
- Give them AI tools that actually work (not low-quality chatbots)
Key Insight: NOC Transformation Is a 5-Year Journey, Not a Lightbulb Moment
The Network Operations Center of the future is fundamentally different from today's NOC. Today's NOC is humans watching dashboards and manually responding to alerts. Tomorrow's NOC is AI watching dashboards and responding to alerts, with humans watching the AI and handling the unexpected.
But you can't make this transition in one quarter. It requires:
- Year 1: Build monitoring and alerting infrastructure that's AI-ready (this alone is a project)
- Year 2: Implement AI-driven alerting and anomaly detection (AI spots problems humans miss)
- Year 3: Implement automated remediation (AI fixes problems without human intervention)
- Year 4: Implement autonomous operations with human oversight (AI is running the NOC, humans are doing incident response)
- Year 5: Mature autonomous operations (the NOC is AI-native, humans focus on strategic work)
During this transition, you need fewer NOC staff, but you need *different* NOC staff: people who understand AI, who can tune anomaly detection, who can design autonomous remediation, who can investigate why the AI made a decision.
Key Insight: Data Teams Must Integrate with Infrastructure Teams
The biggest organizational mistake in AI transformation is keeping data teams separate from infrastructure teams. Data teams have been siloed from infrastructure teams since the dawn of modern IT.
AI obliterates this boundary. To build effective AI systems, you need:
- Infrastructure teams that understand data requirements
- Data teams that understand infrastructure constraints
- A unified data and infrastructure platform
Organizations that are winning AI are reorganizing around "Data Infrastructure" or "Data Platform Engineering", a unified team that owns both the infrastructure and the data flowing through it.
Key Insight: Security, Governance, and Audit Must Converge
Traditional IT splits security (preventing bad things), governance (making decisions), and audit (checking if you followed the rules) into separate teams. AI requires them to work together.
An AI system that has a security incident but passes governance might expose you to risk. A system that passes audit but violates governance might be non-compliant. A system that's secure but nobody understands how it works is a regulatory nightmare.
Forward-thinking organizations are consolidating these functions into a "Risk and Compliance" organization that does security, governance, and audit together.
Practical Use Cases
Use Case 1: The Large Financial Institution Reorganizing the NOC
You're a VP of IT Operations at a large bank. Your NOC has 150 people spread across three shifts. It costs millions per year. They manually review logs, manually respond to alerts, and manually document what happened. The organization's tolerance for outages is extremely low (regulated industry). You need to reduce costs while maintaining or improving reliability.
Your plan:
Year 1: Build an AI-native monitoring stack (one new hire, a senior architect)
- Implement time-series analysis
- Build AI-powered alerting
- Integrate with your existing ticketing system
- Train NOC staff on new tools
Year 2: Implement AI-driven anomaly detection (hire 2 more people)
- AI learns what "normal" looks like
- AI spots anomalies humans miss
- Reduce false positives through tuning
- Begin automated alert enrichment
Year 3: Implement automated remediation (reorganize existing staff)
- AI runs runbooks automatically
- NOC staff move to verification/escalation
- Reduce NOC to 100 people
- Retrain remaining staff
Year 4: Implement autonomous operations (hire 1 governance person)
- AI manages most operational decisions
- Humans handle exceptions
- Reduce NOC to 75 people
- Shift from reactive to predictive
Year 5: Mature autonomous operations (maintain and optimize)
- Reduce NOC to 50 people
- These 50 are now engineers, not operators
- They focus on improving the autonomous systems
- They handle true incidents and strategic improvements
Cost impact: You go from 150 people costing $15M/year to 50 people costing $5.5M/year, while improving reliability and reducing incident response time. The transition costs $2M in infrastructure and tools over 5 years. ROI: positive by year 3.
Use Case 2: The Mid-Market Company Building an AI-Augmented Help Desk
You're an IT Director at a 3,000-person company. Your help desk has 20 people handling 200 tickets per day. Your ASA (answer service level) is 4 hours. You want to improve this without hiring more staff.
Your plan:
Month 1-2: Assess and Plan
- Audit the top 50 ticket types
- Identify which can be answered by AI (typically 30-40%)
- Build the knowledge base
- Design the chatbot flows
Month 3-4: Build the AI infrastructure
- Implement a chatbot platform (your choice: custom LLM, commercial solution, internal)
- Integrate with your ticketing system and ITSM platform
- Start with password resets and common software requests
- Implement human review and feedback loops
Month 5-6: Deploy Phase 1 with training
- Train help desk staff on AI tools and operations
- Deploy chatbot for self-service password resets
- Monitor chatbot performance
- Measure impact (should reduce tickets by 15-20%)
Month 7-9: Expand and retrain
- Add more ticket types to chatbot
- Implement AI-assisted ticket routing
- Begin training help desk on AI-augmented support tools
- Measure impact (should reduce tickets by 30-40% by month 9)
Month 10-12: Scale and optimize
- Help desk focuses on complex issues and customer relationships
- Chatbot and AI-assisted tools handle routine work
- Measure customer satisfaction and staff satisfaction
- Plan year 2 improvements
Impact: Help desk ticket volume drops 35%, ASA drops to 2 hours, help desk staff are happier (they do more interesting work), hiring freezes. No reduction in staff, but redirection to higher-value work.
Use Case 3: The Large Tech Company Reorganizing Around Data Infrastructure
You're a CIO at a large technology company with 200 engineers in Data and 300 engineers in Infrastructure. They build different systems, use different tools, and almost never talk. You're building 40 AI applications, and they're all struggling with data pipeline integration and infrastructure issues.
Your plan:
Phase 1: Structural Change (Q1)
- Create a new "Data Infrastructure" org reporting to VP of Engineering
- Move 50 data engineers and 50 infrastructure engineers into this org
- Hire a VP of Data Infrastructure who understands both sides
- Keep data scientists and analytics in separate "Data Science" org
Phase 2: Unified Platform (Q2-Q3)
- Data Infrastructure team builds a unified data and infrastructure platform
- This platform has data pipelines, feature stores, and infrastructure all integrated
- All 40 AI applications now use this platform
- Consistency, efficiency, and speed improve dramatically
Phase 3: Distributed Specialists (Q4)
- Data scientists and infrastructure specialists embed in business units
- They use the central Data Infrastructure platform
- Application teams move faster because the plumbing is solid
- Data Infrastructure team grows to 150 people but now supports 500+ engineers
Impact: AI application time-to-production drops 60%, infrastructure costs drop 30%, teams move faster, engineers are happier. Net result: same headcount, much better productivity.
Examples
Example 1: Before and After Org Chart - Mid-Market Company
Before:
CIO
├── VP of Infrastructure (60 people)
│ ├── Systems Administration (20)
│ ├── Database Administration (15)
│ ├── Network Operations (15)
│ └── DevOps/SRE (10)
├── VP of Application Support (30 people)
│ ├── Help Desk (20)
│ └── Application Support (10)
├── Chief Data Officer (15 people)
│ ├── Data Engineers (8)
│ └── Data Scientists (7)
└── Director of Security (20 people)
├── Network Security (8)
├── Application Security (7)
└── Audit & Compliance (5)
After (18 months):
CIO
├── VP of AI & Data Infrastructure (40 people)
│ ├── AI Platform Engineering (8)
│ ├── Data Engineering (15)
│ ├── Systems & Infrastructure (10)
│ └── Automation Engineers (7)
├── VP of AI Operations (25 people)
│ ├── AI-Augmented Help Desk (10) [was 20, but more productive]
│ ├── Autonomous NOC (8) [was 15, reduced staff]
│ └── AI Operations Engineers (7) [new role]
├── Chief Data Officer (15 people) [unchanged]
│ ├── Data Scientists (8)
│ └── Analytics Engineers (7)
├── Director of AI Governance (8 people) [new organization]
│ ├── Governance Officer (1)
│ ├── Security Specialists (4)
│ └── Compliance (3)
└── Director of Security (15 people) [reduced]
├── Network Security (8)
├── Application Security (4)
└── [Compliance moved to AI Governance]
Net result: 160 people before, 155 people after. But the 155 people are much more productive. Costs stay flat. Capability grows 40%.
Example 2: Job Description Evolution - NOC to Autonomous Operations
Traditional NOC Operator
- Monitors dashboards and alerts
- Responds to tickets from monitoring system
- Documents resolution
- Average tenure: 2-3 years
- Salary: $55K-$70K
AI-Era NOC Operator (Year 1)
- Monitors AI-generated alerts and remediation
- Reviews AI decisions and escalates exceptions
- Tunes anomaly detection models
- Documents incidents and outcomes
- Average tenure: 3-4 years
- Salary: $70K-$85K
Autonomous Operations Engineer (Year 3)
- Designs automated remediation workflows
- Tunes and improves anomaly detection
- Investigates why the AI made specific decisions
- Improves the autonomous systems
- Average tenure: 4+ years
- Salary: $100K-$130K
Note: You're moving from "operators" to "engineers." Different skill set, different market price, different value creation.
Example 3: Team Composition Evolution - Help Desk
Traditional Help Desk Team (20 people)
- 1 Manager
- 15 Level 1 Support (ticket answerers)
- 3 Level 2 Support (specialists)
- 1 Knowledge Manager
AI-Augmented Help Desk (15 people, after 12 months)
- 1 Manager
- 4 AI Operations Specialists (managing chatbot, AI tools, and complex tickets)
- 6 Specialist Support (handling complex issues, consulting with business units)
- 1 Knowledge Manager
- 2 AI Training Specialists (training staff, tuning models, managing feedback loops)
- 1 Analytics (measuring performance, identifying improvement opportunities)
Note: Net reduction of 5 people, but they're doing more interesting work. Job quality improved, even though headcount decreased. This is what allows the organization to not lay people off, but instead redeploy them.
Example 4: Transition Timeline for Large NOC Reorganization
Timeline
AI Capability
NOC Staffing
Focus
Today
Manual monitoring
150 people, 3 shifts
React to alerts
Month 6
AI-powered alerting
140 people
Learn new tools
Month 12
Anomaly detection
130 people
Tune models
Month 18
Automated remediation
110 people
Verify AI decisions
Month 24
Autonomous operations
75 people
Oversee AI, handle exceptions
Month 36
Mature autonomous ops
50 people
Engineer improvements, incident response
Anti-Patterns
Anti-Pattern 1: "We'll Add AI to Our Existing Structure"
You decide to build an AI help desk chatbot and drop it into your existing help desk structure without changing anything else. What happens: the chatbot answers 10% of tickets, help desk staff don't know how to work with it, they keep doing things the old way, and you have a failed AI project and a frustrated team.
The fix: redesign the team around AI. Change job descriptions. Change KPIs. Change training. Change incentives. Structure change must come with role change.
Anti-Pattern 2: "We'll Just Reduce Headcount"
You implement AI automation and immediately lay off the people who used to do that work. What happens: the remaining team is demoralized, you lose institutional knowledge, your best people leave first, and you're left with people who couldn't leave. Your organization becomes a less attractive place to work.
The fix: use AI-driven productivity gains to redeploy people to higher-value work, not to reduce headcount. This is both more humane and more effective. You keep your best people and redeploy their talent.
Anti-Pattern 3: "Centralized Platform Owns Everything"
You create a central AI Platform team and make every other team go through them to deploy anything AI-related. What happens: everything becomes a bottleneck. Small teams wait months for features. Big teams build around the platform. Velocity goes down. Teams get frustrated and leave.
The fix: build a platform that's good enough that teams want to use it, not a gatekeeper that forces them to. Create hub-and-spoke, not fully centralized.
Anti-Pattern 4: "Keep Everything Separate"
You build an AI Platform team, a help desk AI team, a NOC AI team, and each one builds their own everything. What happens: massive duplication, no consistency, vendors love you, your budget explodes, and you have three different "AI platforms" that can't talk to each other.
The fix: ruthlessly share common infrastructure. Push duplication down to business logic, not to infrastructure.
Anti-Pattern 5: "Data Teams and Infrastructure Teams Are Separate"
You have a Data Engineering team and an Infrastructure team that don't coordinate. Data engineers build pipelines without considering infrastructure. Infrastructure engineers build systems without understanding data requirements. Nothing integrates well.
The fix: create a unified Data Infrastructure org or at minimum, embed data engineers with infrastructure teams and vice versa. Make them collaborate by structure.
Anti-Pattern 6: "The Transition Happens Overnight"
You announce a reorganization, move people around, and expect them to perform immediately. What happens: chaos, mistakes, talent leaves, and the organization spends six months recovering.
The fix: transition gradually. Overlay the new structure on top of the old structure. Let teams work in both modes for a period. Gradually shift responsibility. Give people time to learn new roles.
Human Judgment Checkpoints
Before you reorganize, ask yourself:
Checkpoint 1: Have we clearly defined what success looks like in the new structure? (Not "faster" or "cheaper," but specific metrics.)
Checkpoint 2: Do we have the skills in-house to run the new structure, or do we need to hire? If we need to hire, do we have 6-12 months to find people, or are we reorganizing faster than we can staff?
Checkpoint 3: Have we communicated to the teams what's changing and why? Do people understand it's not about headcount reduction?
Checkpoint 4: Have we identified the people who will struggle in the new structure and planned support for them (retraining, mentoring, or if necessary, outplacement)?
Checkpoint 5: Are we giving people enough time to develop new skills, or are we reorganizing faster than people can learn?
Checkpoint 6: Have we aligned compensation with the new structure? If we're asking help desk staff to do AI work, are we paying them for it?
Executive Summary
Reorganizing IT around AI is not about adding headcount; it's about reshaping how existing headcount creates value. Teams that grow (Platform Engineering, Data Engineering) should be staffed first. Teams that transform (Help Desk, NOC, Database Administration) should transition gradually, with retraining and new tooling happening in parallel. The organizational pattern that works best is hub-and-spoke: a strong central platform team with embedded specialists in operational teams. The biggest mistake is leaving teams siloed (Data separate from Infrastructure, Security separate from Governance). The transition takes 18-36 months, not 90 days. Done well, organizations maintain headcount but increase capability by 30-50% and improve employee satisfaction.
Key Takeaways
Recognize that reorganization for AI is not the same as hiring for AI. You're reshaping your existing organization, not bolting on new pieces.
Map which teams grow (Data Engineering, Platform Engineering), which transform (Help Desk, NOC), and which merge (Data and Infrastructure). This is your roadmap.
Choose an organizational pattern early: centralized platform for small organizations, hub-and-spoke for medium-to-large organizations. Make this choice intentional.
Redesign jobs, not just organizations. Job descriptions, KPIs, career paths, and compensation all need to change to match the new structure.
Integrate teams that have been separate. Data and Infrastructure should work together. Security, Governance, and Audit should work together. The old silos don't work anymore.
Transition gradually. Don't reorganize overnight. Overlay new structures on old ones. Give people time to learn new roles. Support those who struggle.
Communicate relentlessly. Explain what's changing, why it's changing, and what it means for individuals. Address fears about headcount reduction head-on.
Measure the impact. Define success metrics before you reorganize. Measure whether you've achieved them. Be willing to adjust if the structure isn't working.
Remember that organizational structure is how you execute strategy. If your structure isn't aligned with your AI strategy, you won't execute it, no matter how good the strategy is.
Skill.re