1.3: AI Readiness Assessment for Research Organizations
Why Formal Readiness Assessment Matters Before AI Deployment
Research institutions that deploy AI systems without prior readiness assessment consistently encounter two predictable failure modes: wasted investment and faculty alienation. The wasted investment failure happens when an institution purchases or builds AI infrastructure that sits underutilized because the workforce lacks the skills to exploit it, the data governance structures are too immature to supply clean inputs, or the institutional culture treats the initiative as an IT project rather than a research transformation. Faculty alienation occurs when AI systems are introduced without adequate preparation, training, or consultation, producing a backlash that can set an institution's AI trajectory back by years.
Formal readiness assessment prevents both failure modes by establishing an honest evidence base before commitments are made. When a provost or vice president for research can see a dimensional readiness report showing that, for example, infrastructure scores a 4 out of 5 but culture scores a 1.8 out of 5, the investment case for a large GPU cluster weakens considerably. The assessment transforms a political conversation about AI investment into a data-driven conversation about sequenced capability building.
Beyond preventing waste, assessment serves three positive functions. First, it surfaces hidden strengths: departments that have quietly developed sophisticated data pipelines or faculty who have been training models on personal cloud accounts represent assets that institutional leadership often does not know exist. Second, it generates a baseline against which future progress can be measured, which is essential for reporting to boards, accreditors, and funding agencies. Third, it creates organizational momentum: the assessment process itself, which requires interviews, surveys, and cross-functional conversations, often produces more shared understanding of AI's strategic importance than any number of executive presentations.
The appropriate trigger for a formal assessment is any of the following: a strategic plan renewal cycle, receipt of a large AI-related gift or grant, announcement of a competitor institution's AI initiative, regulatory or accreditor guidance touching AI research practices, or a request from the board or senior leadership for an AI strategy. Institutions that wait until they feel fully ready for AI typically wait too long; the assessment is itself a readiness-building exercise.
The AI Readiness Index: Five Dimensions Defined
The AI Readiness Index (ARI) provides a structured framework for evaluating institutional preparedness across five dimensions, each scored on a 1-to-5 scale. The composite ARI score is the average of the five dimension scores, though most strategic conversations focus on the dimensional profile rather than the composite, because a high composite can mask critical weaknesses.
Dimension 1: Infrastructure (scored 1-5). This dimension assesses the institution's physical and digital compute, storage, and networking capabilities. A score of 1 indicates no dedicated AI infrastructure, researchers are using personal laptops or consumer cloud accounts. A score of 2 indicates shared general-purpose HPC clusters without GPU capability. A score of 3 indicates on-premise GPU nodes (typically 8-32 GPUs) accessible to researchers with some friction. A score of 4 indicates a managed GPU cluster with job scheduling, containerization, and a dedicated research computing support team. A score of 5 indicates a comprehensive AI infrastructure stack including on-premise high-performance clusters, integrated cloud burst capacity, InfiniBand networking, high-performance parallel storage, and a self-service portal with model registry and MLOps tooling.
Dimension 2: Talent (scored 1-5). Talent readiness encompasses ML engineers, data scientists, domain experts with AI skills, and research software engineers. A score of 1 means there are essentially no people with AI/ML technical skills employed by the institution outside of one or two CS department faculty. A score of 2 indicates a small number of graduate students or postdocs with ML skills, typically concentrated in one department. A score of 3 indicates the presence of at least one professional data scientist or ML engineer in a centralized research computing or data science hub, with scattered pockets of departmental expertise. A score of 4 reflects a staffed data science center with multiple professional data scientists, ML engineers who can support research projects, and a training program for graduate students. A score of 5 means the institution has a mature talent ecosystem: centralized ML engineering, embedded RSEs in major research centers, a graduate curriculum in AI/ML methods, and active faculty recruitment in AI-adjacent fields.
Dimension 3: Data (scored 1-5). This dimension assesses data quality, accessibility, and governance maturity. A score of 1 means most research data is siloed on individual investigator laptops, poorly documented, and not accessible to others. A score of 2 indicates data is stored in departmental servers with ad hoc organization. A score of 3 reflects a research data management program exists, a data catalog is in place, and FAIR (Findable, Accessible, Interoperable, Reusable) principles are being implemented for some datasets. A score of 4 means the institution has a federated data infrastructure with a searchable catalog, documented data governance policies, data quality standards, and processes for dataset access requests. A score of 5 reflects a mature research data ecosystem: curated datasets available through a governed data commons, automated data quality monitoring, sensitive data enclaves for regulated data, and integration with national data repositories.
Dimension 4: Governance (scored 1-5). Governance readiness covers AI-specific policies, IRB alignment for AI research protocols, and oversight mechanisms. A score of 1 means no AI-specific policies exist and the IRB has no established guidance for AI research. A score of 2 indicates the institution is aware of the need for AI governance and has formed a committee, but no policies have been adopted. A score of 3 reflects adoption of a basic AI use policy (typically covering faculty and student use of AI tools) and initial IRB guidance for AI-involved protocols. A score of 4 means a comprehensive AI governance framework is in place: policies covering research integrity, data use, bias testing, publication standards, and vendor relationships, with IRB expertise specifically in AI-related ethical review. A score of 5 indicates a mature AI governance ecosystem with a standing AI Ethics Board, published algorithmic accountability standards, proactive regulatory monitoring, and external audit mechanisms.
Dimension 5: Culture (scored 1-5). Culture is the most difficult dimension to score and the most consequential for sustainable adoption. A score of 1 means significant faculty skepticism or hostility toward AI, administrative indifference, and no student demand for AI training. A score of 2 indicates some early adopter faculty and student interest, but no institutional commitment or administrative advocacy. A score of 3 reflects visible faculty champions, at least one administrative sponsor at the dean or VPR level, student demand for AI coursework, and positive coverage in internal communications. A score of 4 means AI is a recognized institutional priority with broad faculty engagement, dedicated administrative support structures, and a growing culture of AI-enabled inquiry across disciplines. A score of 5 indicates a pervasive culture of AI innovation: faculty routinely incorporate AI methods, administrators are conversant in AI strategy, students expect AI fluency as part of their education, and the institution is recognized externally as an AI-forward research environment.
Assessment Methodology: Interviews, Surveys, and Audits
A credible ARI assessment draws on three distinct data collection methods, each addressing different aspects of the five dimensions. Relying on a single method, especially a self-reported survey alone, introduces systematic bias and misses crucial qualitative information.
Structured Stakeholder Interviews. The assessment team should conduct 45-to-60-minute interviews with 15 to 25 stakeholders representing different institutional vantage points. Required interview subjects include: the Provost or VPR (to establish strategic context and leadership commitment), the CIO and director of research computing (infrastructure), two to four department chairs from diverse fields (culture and talent), the IRB chair and research compliance officer (governance), the chief data officer or institutional research director (data), and four to eight faculty members stratified by seniority, discipline, and prior AI experience. Each interview follows a semi-structured protocol with questions calibrated to the interviewee's domain, but all interviews include three universal questions: What AI-related activities are already happening in your area? What would need to be true for AI research to flourish here? What is the single biggest barrier you see?
The 20-Question Faculty Survey. The faculty survey targets a sample of at least 40 to 60 faculty across departments and ranks (assistant, associate, full, emeritus). The 20 questions are organized into five blocks of four, one block per ARI dimension, and use a mix of Likert scales (strongly agree to strongly disagree), multiple-choice capability questions, and a small number of open-ended items. Key questions include: frequency of AI/ML tool use in current research; perception of data quality for AI research purposes; awareness of institutional AI policies; perception of administrative support for AI; and willingness to participate in AI training programs. Survey anonymity is essential for honest responses, particularly on the Culture dimension.
IT Infrastructure Audit. The research computing team should provide a technical audit document covering: total GPU count and GPU-hours available per month, storage systems and aggregate capacity, network bandwidth and latency between compute and storage, current job scheduler and container runtime, active software licenses, and help desk ticket volume and resolution time for AI-related requests. The assessment team cross-validates self-reported infrastructure capability against actual utilization data: a cluster that is nominally available but runs at 95% utilization with two-week queue times scores lower than one that runs at 60% utilization with same-day access.
Benchmarking Against EDUCAUSE Data. The EDUCAUSE Annual Higher Education Technology Report and the AI in Higher Education study provide comparative data from hundreds of peer institutions. The assessment team should identify a cohort of 10 to 15 peer institutions (matched on Carnegie classification, research expenditure level, and enrollment) and benchmark the subject institution's infrastructure, talent, and governance practices against this peer cohort. Benchmark positioning (e.g., "our GPU-per-researcher ratio is at the 30th percentile of R2 peers") is more persuasive to leadership and boards than absolute numbers because it frames investment as a competitive necessity rather than an aspirational luxury.
The Readiness Report: Format and Framing
The readiness assessment report is a strategic document, not a technical audit. Its primary audience is institutional leadership, provost, VPR, CFO, board, who will use it to make resource allocation decisions. A well-structured report follows a consistent format that makes findings accessible and actionable.
Executive Summary (2 pages maximum). The executive summary states the assessment scope and methodology, presents the five dimension scores in a visual radar chart or spider diagram, identifies the two to three most consequential gaps, and states a recommended path forward. Busy leaders read the executive summary first and sometimes exclusively; it must be self-contained and free of jargon.
Dimensional Scorecards. Each of the five dimensions receives its own section presenting: the dimension score with scoring rationale (citing specific evidence from interviews, survey results, and audits), a comparison to peer institution benchmarks, a narrative description of current strengths and specific gaps, and a set of two to four targeted recommendations for improvement. The scoring rationale must be specific: "Culture scores 2.3 because 68% of survey respondents reported they do not feel supported by their department chair in exploring AI methods (Q14), and only 3 of 22 interview subjects could name a concrete institutional AI initiative (qualitative theme 4)." This specificity builds credibility with skeptical audiences.
Gap Analysis Prioritized by Impact and Feasibility. The gap analysis synthesizes findings across all five dimensions to identify the most critical capability deficits. Each gap is characterized along two axes: impact (what would be gained if this gap were closed, expressed in terms of research outcomes, faculty recruitment, and grant competitiveness) and feasibility (how difficult is closure given available resources and change management requirements). High-impact, high-feasibility gaps are immediate priorities. High-impact, low-feasibility gaps require phased multi-year strategies. Low-impact gaps of any feasibility may be deferred.
The 6-Month Action Plan with Named Owners. Recommendations without owners are wishes. The action plan section converts assessment findings into concrete 90-day and 6-month actions, each assigned to a named institutional role (not a committee, a specific person). A well-formed action item specifies: what will be done, who is responsible, what resources are required, what the measurable success indicator is, and by what date. Typical 6-month action plan items include: forming an AI governance working group (owner: VPR, by month 2), launching a faculty AI interest survey to identify potential champions (owner: research development director, by month 1), issuing a Request for Information for on-premise GPU cluster vendors (owner: CIO, by month 3), and submitting an NSF CC* infrastructure proposal (owner: research computing director, by month 6).
Common Readiness Assessment Pitfalls
Experienced institutional researchers and AI strategy consultants have documented a consistent set of assessment errors that undermine the validity and utility of readiness findings. Awareness of these pitfalls does not make them disappear, but it allows assessment teams to design protocols that actively mitigate them.
Overestimating Infrastructure. The most common infrastructure assessment error is conflating hardware existence with hardware accessibility. An institution may own 128 GPUs that are technically available for research, but if those GPUs are managed by a single department and are not accessible through a shared scheduling system, their effective capacity for institutional AI research is close to zero. The audit must distinguish between nominal infrastructure and effective infrastructure, what researchers can actually use without exceptional permissions, political capital, or workarounds. A secondary overestimation error occurs when institutions count cloud access as infrastructure readiness without accounting for the friction of procurement, the lack of training in cloud ML services, and the absence of cost management controls that routinely cause researchers to exceed budgets and lose cloud access.
Underestimating Culture Change. Research institutions have strong disciplinary cultures with deep roots in specific epistemological traditions. A faculty cohort of quantitative social scientists may embrace AI methods readily; a cohort of humanists or qualitative ethnographers may resist them for principled intellectual reasons rather than mere unfamiliarity. Assessment teams frequently assign culture scores based on the presence of faculty champions and administrative rhetoric without adequately probing the distribution of attitudes across the full faculty. A single AI-enthusiastic VPR can produce the illusion of culture readiness while the median faculty attitude is indifference or active skepticism. The assessment must measure both the mean and the variance of faculty attitudes.
Conflating Data Availability with Data Usability. Institutions often cite the existence of large institutional datasets, student records, research output databases, biological sample libraries, longitudinal survey archives, as evidence of data readiness. But availability is not usability. Data usability for AI research requires: consistent schemas across time periods, documented provenance, accessible formats (not locked in proprietary systems), and governance clearance for the intended research uses. A dataset that exists but is locked in a 20-year-old enterprise system with no API, inconsistent variable names across years, and no data dictionary is essentially unavailable for AI research regardless of its nominal size. Assessment teams should audit a sample of frequently cited datasets for actual usability and report usability scores separately from availability scores.
Anchoring on Best-Case Scenarios. Readiness assessments conducted entirely or primarily by internal staff, without external benchmarking or an external validator, tend to produce optimistic scores. Internal assessors are subject to social pressure to report favorably on colleagues and to institutional pride that distorts self-evaluation. The best practice is to include at least one external reviewer (a peer institution CRIO, an AI-experienced board member, or an external consultant) who validates the scoring rationale and provides an independent assessment of whether the score accurately reflects observed evidence.
Benchmark Targets by Institution Type
AI readiness expectations and target scores should be calibrated to institutional type. An R1 research university competing for NIH Big Data to Knowledge (BD2K) awards and NSF AI Institutes has different readiness requirements than a primarily undergraduate institution (PUI) or a regional R2 university. Setting uniform readiness targets across institution types produces either premature ambition (for smaller institutions) or false comfort (for large research universities whose resources should generate higher scores).
R1 Research Universities. R1 institutions receive the majority of their revenue from sponsored research and compete in national and international markets for elite faculty and graduate students. For these institutions, infrastructure target scores should be 4.0 or higher, reflecting the expectation of a managed GPU cluster with self-service access and MLOps tooling. Talent scores should target 4.0, requiring a staffed data science center and embedded RSE capacity. Data scores should target 3.5, reflecting the complexity of managing diverse research datasets across many disciplines. Governance should target 3.5, and Culture should target 3.0 as a realistic medium-term goal given disciplinary diversity. An R1 with a composite ARI below 3.0 should treat that as a serious strategic vulnerability.
R2 Research Universities. R2 institutions have substantial research activity but smaller sponsored research portfolios. Infrastructure targets should be 3.0, representing a shared GPU cluster that may require some advance scheduling. Talent targets should be 3.0 (a data science hub with professional staff). Data and Governance targets should be 3.0. Culture targets should be 2.5, reflecting realistic expectations for a smaller faculty. A composite ARI of 2.5 to 3.5 is a reasonable 3-year horizon for most R2 institutions.
Primarily Undergraduate Institutions. PUIs face AI readiness challenges that are fundamentally different in kind rather than merely smaller in scale. Faculty at PUIs have heavy teaching loads that leave little time for research tool adoption. The student audience (undergraduates) creates different training needs than a graduate-student research workforce. Infrastructure targets for PUIs should be 2.5, reflecting shared access to cloud compute (AWS SageMaker, Google Colab Pro) rather than on-premise clusters. Talent targets should be 2.0, recognizing that PUIs typically do not have and should not expect to have dedicated ML engineering staff; instead, one or two computationally skilled faculty serve as local AI champions. The most important readiness dimension for PUIs is Culture (target 3.0), because faculty enthusiasm and student demand are the primary drivers of AI adoption at institutions with limited infrastructure and no graduate research workforce.
Gap Analysis Framework: From Current State to Critical Path
The gap analysis converts dimensional scores into an investment and change management roadmap. It is the operational heart of the assessment report, the section that translates diagnosis into prescription. A rigorous gap analysis follows a four-step framework.
Step 1: Define Current State. For each gap (a specific capability deficit within a dimension), describe current state with precision. "Our data infrastructure is inadequate" is not a useful current state description. "Research data for 78% of active NIH-funded projects is stored on PI-managed servers without backup, not cataloged in any institutional system, and inaccessible to collaborators without ad hoc permission from the PI" is a useful current state description because it can be verified, communicated, and measured against.
Step 2: Define Target State. The target state should be specific, time-bounded, and tied to a strategic outcome. "We will have a research data management platform that catalogs 100% of externally funded research datasets, provides role-based access control, and is compliant with NIH data management and sharing requirements within 24 months." Linking the target to a funding agency requirement (in this case, NIH's 2023 data management and sharing policy) is strategically powerful because it transforms the gap from an internal preference into a compliance necessity.
Step 3: Map the Critical Path. The critical path identifies the sequence of dependencies that must be resolved to move from current state to target state. For the data management example, the critical path might be: (a) select and procure a research data management platform (months 1-4), (b) configure institutional Single Sign-On integration (months 3-5), (c) develop a data deposit workflow and train research administrators (months 5-8), (d) mandate deposit of new grant datasets at time of award (month 9), (e) backfill existing active projects (months 9-24). Each step in the critical path should identify the responsible role, estimated effort, and any dependencies on other critical path steps.
Step 4: Quantify Resource Requirements. Every gap in the critical path analysis must be associated with estimated resource requirements: personnel (FTE or partial FTE over the planning horizon), technology (one-time license/hardware costs and ongoing operational costs), and change management (training, communications, policy development). These estimates feed directly into the budget case. The gap analysis should explicitly connect each resource requirement to a potential funding source: internal reallocation, philanthropic gift, federal infrastructure grant (NSF CC*, NIH S10), or state system funding.
Building the Budget Case from Readiness Findings
One of the most practical functions of the readiness assessment is to generate the evidentiary foundation for budget requests. Research leaders who walk into a CFO or provost meeting with a gap analysis, peer benchmark comparisons, and a prioritized investment plan with estimated costs and funding source options are dramatically more likely to secure resources than those who present abstract AI strategies or enthusiasm-driven proposals.
The budget case structure follows the logic of the gap analysis: gap identification (with evidence), investment required to close the gap (with cost estimates), expected return (expressed in terms of grant competitiveness, faculty recruitment, research output, and indirect cost recovery), and alternative funding sources that reduce the institutional burden. Each gap-to-investment connection should be as specific as possible. Rather than "we need to invest in AI talent," the budget case should specify: "Hiring two Research Software Engineers with ML specialization at $95,000-$110,000 per year each (plus 28% fringe) would provide embedded support to our top 12 AI-active PIs, directly enabling five grant proposals in fiscal year 2027 that represent an estimated $8.2M in direct costs and $3.1M in indirect cost recovery at our current F&A rate."
The budget case should also include a cost-of-inaction analysis: what will the institution lose or forgo if the gap is not closed? Peer institution benchmarks are particularly powerful here. If 80% of R2 peer institutions have a managed GPU cluster and the institution does not, the budget case can credibly argue that the absence of infrastructure is already costing faculty grant competitiveness, particularly for grants with data-intensive computation requirements.
Finally, the budget case should present a phased investment scenario in addition to an ideal-scenario request. Leaders who cannot secure full funding in a single year need a credible Phase 1 that delivers meaningful value and creates momentum for Phase 2 funding. A phased scenario also demonstrates the strategic maturity of the assessment team and builds trust with CFOs who are conditioned to distrust all-or-nothing proposals.
Continuous Readiness Monitoring: Quarterly, Annual, and Trigger-Based
The initial readiness assessment establishes a baseline, but AI technology and institutional capabilities both evolve rapidly. A single assessment quickly becomes stale. Institutions need a continuous monitoring protocol that maintains current situational awareness without imposing the full burden of a comprehensive assessment on a frequent basis.
Quarterly Pulse Assessments. Every quarter, the research computing or VPR's office conducts a lightweight 10-question pulse survey with a rotating panel of 20 to 30 faculty. The pulse survey focuses on three or four indicators in the dimensions experiencing the most active change (typically Infrastructure and Culture during the first two years of an AI initiative). It takes respondents under 5 minutes to complete. Results are compiled into a one-page dashboard that tracks indicator movement relative to the baseline and provides early warning of emerging problems (e.g., a decline in faculty satisfaction with GPU access response times as demand grows).
Annual Full Assessment. Once per year, typically aligned with the strategic planning or budget cycle, the institution conducts a full ARI assessment update. This does not require the full scope of the initial assessment, the interview list can be shorter and more targeted, and the audit can focus on changed components, but it produces an updated dimensional scorecard that tracks progress against the targets set in the gap analysis. Annual updates allow the institution to report AI readiness improvement to its board, accreditors, and stakeholders.
Trigger-Based Reassessments. Certain events should automatically trigger an unscheduled assessment update: receipt of a major grant or gift that will significantly change infrastructure or talent; loss of a key AI leader (chief data scientist, research computing director); a significant cybersecurity incident affecting research data; new federal regulations or funding agency requirements touching AI; or a major technology shift (e.g., the emergence of a new model architecture that changes compute requirements). Trigger-based reassessments are typically narrower in scope than full annual assessments, focusing on the dimensions most affected by the triggering event.
Operationalizing Assessment: Teams, Timelines, and Tools
A practical readiness assessment at a research institution of moderate complexity, 20 to 50 AI-active faculty, 3 to 8 departments, a research computing team of 2 to 5 people, requires approximately 8 to 12 weeks from launch to final report. The assessment team should include a project lead (typically the research computing director or a research development officer with strategic experience), a data collection coordinator, a technical liaison from IT, and an external validator. Total internal effort is approximately 120 to 180 staff-hours spread across the team.
Several commercial and open-source tools support the assessment process. For surveys, REDCap (widely available at academic medical centers) and Qualtrics both provide the anonymization and routing logic needed for the faculty survey. For data management auditing, institutional data catalogs (Collibra, Atlan, DataHub) can provide automated inventories of registered datasets, though most institutions will require significant manual supplementation. For benchmarking, EDUCAUSE's Core Data Service provides benchmarking data for technology staffing and infrastructure; the Association of Research Libraries (ARL) provides data relevant to data services and digital scholarship programs; and the Carnegie Classification provides the peer institution comparison set.
One frequently underestimated operational challenge is interview scheduling. Senior administrators and faculty are busy; a 45-minute interview request often requires 2 to 3 weeks of lead time. The assessment team should send interview invitations before any other assessment activity begins, and should prepare a one-page briefing document explaining the assessment purpose and anticipated time commitment that can be sent with the invitation. Framing the assessment as a strategic opportunity rather than an audit or evaluation reduces scheduling resistance and encourages candor from interview subjects.
Finally, the readiness report should be positioned carefully in the institutional communication sequence. The report itself is an internal strategic document and should be treated as confidential during the leadership review phase. Sharing draft findings too broadly before leadership has had a chance to react, contextualize, and develop a response plan can create anxiety or premature public commitments. Once leadership has reviewed findings and the action plan has been developed, a version of the findings, appropriately summarized and stripped of individual-level attribution, can be shared with faculty governance to demonstrate institutional transparency and build buy-in for the resulting initiatives.
Skill.re