The pace of federal AI policy has been genuinely unusual. Multiple executive orders, OMB memoranda, agency-specific implementation plans, and congressional activity have layered obligations on top of each other faster than most agencies’ acquisition and IT processes can absorb.
For agency leaders trying to figure out what they actually have to do — as opposed to what they’re supposed to aspire to — the signal-to-noise ratio is challenging.
The Operational Requirements That Matter
Cutting through the policy language, a few categories of concrete requirements are driving most of the implementation work.
AI use case inventories. OMB M-24-10 required agencies to publish inventories of their AI use cases by December 2024. This sounds administrative, but doing it right requires someone in each agency to actually know what AI is being used, where, and for what purpose — including shadow deployments that procurement never formally approved. The agencies that treated this as a real accountability exercise learned things about their own AI footprint that surprised them.
Chief AI Officer designations. Agencies are required to designate a Chief AI Officer with specific responsibilities for AI governance, risk management, and coordination. In practice, the CAIO role is being stood up with very different authority levels and resource allocations across agencies. Some have genuine mandates and staff; others are a designation layer over existing responsibilities with no new capacity.
Rights-impacting and safety-impacting AI. The most consequential operational requirement is around AI systems that affect individual rights or safety. These include systems used for benefits eligibility determinations, law enforcement screening, housing assistance decisions, and similar high-stakes uses. The policy requires specific impact assessments, human oversight mechanisms, and — in some cases — the ability to opt out of automated decision-making. Implementing these requirements on legacy systems that were never designed with these constraints is expensive and time-consuming.
Minimum risk management practices. Based on the NIST AI Risk Management Framework, agencies are expected to have consistent processes for evaluating AI systems before deployment, monitoring them in production, and handling incidents. Most agencies had inconsistent or no formal practices before these requirements landed.
Where Agencies Are Struggling
The gap between policy requirement and operational reality is visible in a few specific areas.
Procurement tooling hasn’t caught up. Buying AI responsibly requires evaluation criteria that standard IT procurement processes don’t include. How do you evaluate a vendor’s claims about model accuracy, bias testing, or data provenance through FAR-based solicitation processes? The guidance is improving, but contracting officers are often working without clear precedents.
Workforce capacity is thin. The implementation work requires people who understand both AI technology and the specific mission context well enough to make good decisions. That intersection is rare in government. Training programs are improving but can’t manufacture expertise overnight.
Legacy system integration is harder than anticipated. The most critical AI use cases — the ones affecting real people — are often embedded in or adjacent to legacy systems that are decades old. Retrofitting human oversight mechanisms, audit logging, and bias monitoring onto systems that weren’t designed for any of this is genuinely hard technical work.
The “minimum viable compliance” risk. Agencies under resource pressure are tempted to check the policy boxes with minimal actual implementation. Publishing an AI use case inventory with vague entries. Designating a CAIO with no staff. Producing impact assessments that don’t reflect real analysis. This approach creates paper compliance and real risk — and the oversight community is getting better at telling the difference.
What Good Implementation Looks Like
The agencies making real progress share a few characteristics. They started with honest inventories of what AI they were actually using. They designated CAIOs with real authority and connected them to the acquisition process. They prioritized the highest-stakes use cases for the most rigorous oversight requirements instead of trying to apply everything everywhere simultaneously.
They also treated the NIST AI RMF as a practical tool rather than a compliance framework — using the Govern, Map, Measure, and Manage functions to build processes that work for their specific context, not just check boxes.
The trajectory of federal AI governance is toward more accountability, not less. Agencies that build genuine capacity now will be better positioned than ones that optimize for the current minimum requirement and have to rebuild when the bar moves.
If you’re building out your agency’s AI governance program and looking for practical frameworks rather than compliance theater, that’s a conversation worth having.