Much of the federal policy focus on AI is currently directed at frontier model evaluation. Sitting quietly next to it is another directive that asks federal agencies to strengthen the operational defenses behind citizen services. That second directive is where the more significant operational challenge lies. And it isn’t a shortage of tools. It’s that many agencies still struggle to get the ones they’ve already paid for to work together.
In many federal agency operations teams, somewhere between eight and fifteen platforms are in active daily use. Consoles for monitoring, logging, ticketing, vulnerability management, change approval, runbook automation and secrets management—plus a few more added after the last big incident. Each was bought for a good reason. Each has its own data model and its own people trained to run it.
The IBM X-Force 2026 Threat Intelligence Index reports public-facing application exploits increased 44% year over year. The workforce defending that attack surface is also responsible for maintaining every console. The result is a gap between what the systems see and what teams can act on. The site reliability engineer (SRE) sees one screen, the security analyst another and the compliance officer a third. Meanwhile, the CFO is defending the operations budget to Congress. Every view is correct, but none is complete.
The person in the middle has become the integration layer, gathering evidence across half a dozen tools, reconstructing the timeline by hand, and writing a root-cause report that looks an awful lot like the one they wrote three months ago for what was essentially the same incident.
A tool shortage isn’t the problem. Agencies are short on coherence. The reflex to solve that by buying another category leader is part of what got us here. What government needs is one layer up, something that sits across the existing investments and gives the people running the agency’s services one coherent surface to work from. The honest analogy, after years of building these systems, is a coordination layer—an operating system for the operations function itself. This matters more in government than in industry for three reasons:
1. Agencies can’t tear out what they already have.
ServiceNow, Splunk, the monitoring stacks, the SOAR runbooks—agencies have paid for, integrated, accredited and trained staff to use these. Retiring these tools is a multi-year program that federal customers can’t run while keeping mission live. So, the architecture must be agnostic across tools, clouds and models from day one. It needs to sit behind the ticketing system, on top of the telemetry, on whichever sovereign cloud has been accredited, with whichever AI models agency policy allows. In practice, the first agencies adopting this pattern are standing it up on AWS GovCloud—by using Amazon Bedrock for model routing and the surrounding AWS security and audit services to meet FedRAMP High and IL5 obligations from the outset. The same coordination layer runs on any accredited sovereign cloud; AWS is simply where the earliest federal proof points are live today.
2. The federal government cannot treat AI as a black box.
Every decision a federal agent makes—human or machine—must be defensible to an IG, a FedRAMP assessor or a Hill oversight committee. Every AI agent decision should be logged with its inputs and outputs. Every approval should be tied to a named human reviewer. Every remediation should produce a tamper-evident record. This can’t be bolted on later. It must be how the platform was built.
3. The workforce math doesn’t allow waiting.
Hiring constraints, retirements, the difficulty of clearing engineers. Agencies face a structural choice. Either operations work scales differently or the services that people depend on absorb the cost.
When a coordination layer is in place, four very different roles in an agency finally work from the same picture. The reliability engineer opens an incident and finds root cause analysis and a recommended fix already attached. Right alongside, the security analyst sees the vulnerability context and the path to remediation. Compliance teams no longer need to reconstruct the audit trail the next morning because the record is captured as work happens. And the CFO, for the first time in a long time, watches the operations cost curve start to bend down rather than up, as the agents absorb more of the routine work.
For the people on call, this changes what a 3 AM page feels like. Today, a page kicks off a manual scramble. Now the page leads to a workspace where the initial investigative work is already done. Evidence is assembled, a recommendation is on screen and the human’s job is to approve, adjust or escalate. Operations engineers still play a central role, but this model allows teams to scale their workloads without the same level of burnout.
Federal CFOs are starting a conversation almost no one was having eighteen months ago. AI inference—the tokens models process every time they read evidence and write an answer—has become a monthly budget line item. As more operational work shifts from human-driven to AI-augmented, those tokens compound into a bill that scales with mission.
The discipline this calls for is what many teams now describe as a token economy. Not every task needs the most expensive AI available. Triaging alerts, classifying tickets, summarizing runbooks. Smaller, lower-cost models handle these just fine. Where premium frontier models earn their cost is in the hard stuff: root-cause reasoning across distributed evidence such as the 4 AM cascading incident where a more advanced model adds measurable value. The platform should route those calls automatically, matching the size of the intelligence to the size of the problem.
And the cost difference isn’t small. Across federal workloads we’ve looked at, running everything on a top-tier model versus routing to right-sized models can raise monthly costs by several multiples for comparable operational outcomes. That gap often decides whether an AI strategy survives a budget review. It also gives the model-agnostic argument a CFO edge. An agency locked into one provider is locked into one price curve.
AI-enabled agentic operations is positioned to become a standard layer in the federal IT stack—the same way two-factor identity and continuous monitoring became standard a decade ago. The real decision this year is whether to keep buying tools or finally bring in the layer that can improve coordination across existing tools. Agencies that move early may be able to support the same mission with the same headcount, handle several times the workload and present their CFOs a more defensible cost profile.
Meet IBM at AWS Summits in Washington, DC from 30 June – 1 July
Turn your application data into actionable insights, helping you strengthen operations and improve IT resilience.
End-to-end visibility and intelligent insights to help you detect issues early, reduce downtime and keep your applications resilient.
Detect issues early, predict outages and help you prevent disruptions before they impact your business.