Table of Contents

A state agency pilots an AI tool to flag likely improper payments in a benefits program. In the demo, on a clean sample dataset, it works. In production, it reads from the same case management system caseworkers use every day. It flags the wrong cases, misses obvious ones, and nobody can explain why.

The model didn’t change between the demo and production. The data it had to reason over did.

That pattern shows up across government. In NASCIO’s 2026 State CIO Survey, 82% of states reported AI projects in production. Only 35% reported AI running at enterprise scale across the executive branch. What separates the two is data readiness for AI: whether the data underneath was ever made ready for a machine to act on.

What Is Data Readiness for AI?

For most of the last decade, data readiness meant proximity to a dashboard. If data was stored, modeled, and visualized, an analyst could make sense of it. That bar assumed a human was always in the loop to fill in what the system left out.

AI removes that human. The Modern Data Report 2026 found that models don’t browse Slack channels or remember why one metric is “more correct” than another. An AI agent, a system that carries out multi-step tasks and calls data sources on its own, only knows what the system makes machine-readable.

Infographic depicting results from The Modern Data Report 2026, suggesting how metrics from different sources and domains conflict each other in enterprises, where 93% cite how they encounter conflicting versions of the same metric. | The Modern Data Company
The illusion of alignment in metrics | Source: The Modern Data Report 2026

The report surveyed more than 540 data leaders and practitioners across 64 countries. Ninety-three percent report encountering conflicting metrics. Sixty-eight percent say their data isn’t trustworthy enough for AI to act on directly.

Animesh Kumar, Co-Founder and CTO of The Modern Data Company, draws a clear line between analytics-ready and AI-ready data. AI-ready data passes three tests:

  • Fresh: current enough that a model isn’t reasoning about a world that no longer exists.
  • Complete: granular enough to avoid brittle, logically incomplete outputs.
  • Semantically rich: relationships, hierarchies, and intent are encoded in the data itself, not left for a human to infer.
A table view of how analytics-ready data is different from AI-ready data in aspects like freshness of data, completeness of data, and readability of data by human and machine consumers. | The Modern Data Company
Analytics-ready data vs. AI-ready data

Data readiness for AI means the data explains itself, because no human will be there to explain it.

Why Is Data Readiness Harder for the Public Sector?

Every organization chasing AI readiness deals with legacy systems and fragmented ownership. Government carries an extra requirement. Agencies serve citizens and every automated decision is subject to oversight.

A review of public sector data stacks by Modern Data 101, the practitioner community supported by The Modern Data Company, names three requirements that follow directly from that distinction:

  1. Auditability. The public, inspectors general, auditors, legislators, and courts need to examine a decision rather than take it on faith. Under OMB M-25-21, high-impact federal AI uses also require documented risk management practices.
  2. Interoperability. Agencies routinely share data across programs and jurisdictions without surrendering control of it.
  3. Adaptability. A policy change has to become a system change without a contract renegotiation in between.

The data itself raises the stakes further. Criminal justice information (CJIS), protected health information (HIPAA, 42 CFR Part 2), federal tax information (IRS Publication 1075), and student records (FERPA) each carry their own access rules. Records retention and FOIA obligations extend to what AI systems produce.

The image shows an overview of the unique challenges of enabling AI at scale in public sector, including different end consumers, policy agility, and jurisdictional sharing. | The Modern Data Company
Challenges of public sector AI

None of that is achievable by adding a compliance layer on top of an unready data estate. Most states have not built that foundation yet. In a NASCIO and EY study of 46 states, only 22% had a dedicated data quality program. Seventy-two percent described their approach to data quality as reactive or merely aware.

Federal agencies show the same gap. In a March 2026 Market Connections survey of more than 200 federal IT executives, nearly 90% said they require logging and audit trails for agentic AI. Fewer than a third had implemented the oversight frameworks most of them called essential.

The Policy Context

Recent guidance makes data readiness a compliance question, not only a technical one:

Data ownership belongs here too. An agency that can’t trace its data, or take it along when a contract ends, can’t prove auditability or interoperability.

In government, data readiness for AI is not a technical preference. It is the evidence an agency needs to defend every automated decision.

The Seven Stages of AI-Ready Data Architecture

Kumar and Travis Thompson, Chief Architect and Head of Engineering at The Modern Data Company, lay out AI readiness as a sequence, not a checklist. Each stage fixes the specific failure the previous stage left behind. Skip one, and the debt lands on the team that hits the wall two stages later, usually analytics and AI.

StageWhat it establishesWhat breaks without it
1. Modular infrastructureStorage, compute, and orchestration as swappable partsEvery change touches everything; no room to add new tools
2. Central + distributed ownershipDomain teams own the data they understand bestA central team becomes a permanent bottleneck
3. Unified data layerOne consistent interface across existing warehouses and lakesAgents learn a dozen different schemas instead of one
4. Governed, observable platformLineage, quality checks, and access policy enforced at the point of useSmall quality issues become organization-wide ones
5. Semantic, context-aware layerOne canonical definition per metric and entityTwo governed, high-quality datasets still disagree on what “active case” means
6. Self-service developer platformTrusted data products any team can discover and reuseOnly the platform team can safely build anything, so nothing scales
7. AI-native platformData products callable directly by agents, policy enforced at runtimeAgents inherit every unresolved ambiguity beneath them, at machine speed

Each stage is a prerequisite for the next, so an agency is only as AI-ready as the lowest stage it has finished.

What Happens When You Skip a Stage?

At stage four, governance stops bad data from moving. It says nothing about whether two clean, well-governed systems agree on what a metric means. That is stage five’s job. A semantic layer, one shared set of business definitions that every tool and agent reads from, settles what “active case” means once.

Gartner puts a number on the gap. It predicts that by 2027, organizations that prioritize semantics in AI-ready data will increase agentic AI accuracy by up to 80% and reduce costs by up to 60%. The seven-stage sequence explains why that gain depends on what sits underneath. Semantics only pays off once governance is in place.

The cost of skipping stages shows up in plainer terms too. The Modern Data Report found that 89% of practitioners rank finding the right data among their top three-time drains. Seventy percent report rework within a single quarter.

Infographic depicting results from The Modern Data Report 2026, suggesting how the discovery stage takes most time, as cited by almost 90%, 74% cited validating trustworthiness and data quality, and 71% cited it takes most time to understand the data | The Modern Data Company
The cost of discovery | Source: The Modern Data Report 2026

None of that is a tooling gap. It’s stages four and five, skipped, charging interest.

How DataOS Enables the AI-ready Data Architecture

DataOS® for Public Sector treats the seven stages as one continuous architecture rather than seven separate purchases. It is built around data products, governed, semantically defined datasets built for a specific use, with lineage and access policy attached. A metric defined once reads the same way to an analyst, a dashboard, or an AI agent.

Treating context as infrastructure lets an agency add a new jurisdiction’s data or a new oversight requirement without rebuilding stages one through four. Lineage, the record of where data came from and how it changed, travels with each data product. So does its access policy. That is what makes auditability and interoperability practical, as the data product ecosystem grows.

For agencies, DataOS works with existing systems rather than replacing them. That matters when budgets run on annual cycles and legacy systems can’t be switched off. DataOS is available through Carahsoft, NASA SEWP V, ITES-SW2, and NASPO ValuePoint, DIR-CPO-5687 and supports agency FedRAMP, TX-RAMP Level 2, CJIS, and NIST SP 800-171 requirements.

An infographic depicting how DataOS enables context as infrastructure by establishing the seven stages of AI-readiness as one continuous out-of-the-box architecture, with data products carrying their own governance, lineage, and semantic definition as part of what it is. | The Modern Data Company
Setting up context as infrastructure with DataOS®

How to Assess Your Own Data Readiness for AI

Before the next AI pilot gets scheduled, find out which of the seven stages your agency is standing on. Most AI accuracy and governance problems trace back to data that was never made ready. The fix starts below the model, not inside it.

Start with seven questions. A “no” marks the stage to fix first.

  1. Modular infrastructure: Can we add or replace a storage, compute, or analytics tool without rebuilding what sits around it?
  2. Ownership: Do program offices own and publish their data, or does every request go through one central team?
  3. Unified access: Can an analyst or AI tool reach data across our warehouses, lakes, and legacy systems through one consistent interface?
  4. Governance: If an inspector general asked which data informed an automated decision, could we show the lineage, quality checks, and access policy in under a day?
  5. Semantics: Do our programs share one definition for terms like “active case,” “eligible household,” or “resolved incident”?
  6. Self-service: Can a program team find, trust, and reuse an existing data product without opening a ticket?
  7. AI-native: Can an AI agent call a governed data product directly, with access policy enforced at the moment it asks?

To score your answers against a structured benchmark, take the AI Readiness Scorecard.

The agencies that scale AI won’t be the ones with the best models. They will be the ones whose data could answer an inspector general before the model ever ran.

Topics: 
Curious how to make AI more reliable in your organization?
Cover of The Modern Data Report 2026 titled The Data Activation Gap with abstract blue and red gradient background.
Get the Report
Find out what your peers are saying.

Continue reading

How DataOS Delivers on NASCIO’s Top 10 State CIO Priorities
Public Sector

How DataOS Delivers on NASCIO’s Top 10 State CIO Priorities

The Modern Data Company
Oct 7, 2026
What The Modern Data Report’s 5 Emerging Trends Mean for the Public Sector
Public Sector

What The Modern Data Report’s 5 Emerging Trends Mean for the Public Sector

Rick Rosenburg
Sep 16, 2026
The Modern Data Company Launches DataOS for Public Sector at the State of Texas DIR Innovation Lab for Demonstrations
Announcement

The Modern Data Company Launches DataOS for Public Sector at the State of Texas DIR Innovation Lab for Demonstrations

The Modern Data Company
Sep 15, 2026
The Modern Data Company Achieves TX-RAMP Level 2 Certification for DataOS
Announcement

The Modern Data Company Achieves TX-RAMP Level 2 Certification for DataOS

The Modern Data Company
Aug 3, 2026
See how DataOS can put data to work for you
Get started →