In 2013, my family had to fight a disability claim on behalf of my father, a disabled veteran, and I’ve managed his care ever since. Back then, the process was manual and it took years: delay after delay, denial after denial, most of them simple mistakes from the people answering the phones. We finally hired a lawyer, who took a cut of benefits he was already owed. The part that still stays with me is that the information proving his eligibility was sitting in a database inside the same system that kept denying him. The answer was there the whole time. No one on the other end of the line owned the job of finding it.

I don’t share that to relitigate the Department of Veterans Affairs, which, like much of the federal government, has worked hard to modernize in the years since. I share it because it taught me something I’ve watched hold true across a career in customer experience analytics, in and out of government: Most service systems, federal ones included, are built to do one job well, and it isn’t solving your problem. It’s ending your call. You reach a federal agency, an agent verifies your information, tells you what’s wrong and what to do next, and moves you along. More often than not, you call back because the information was incomplete or you were routed to the wrong place the first time. The goal of the interaction is speed. It is not resolution.

I’ve named the missing piece “call ownership.” When no one owns a problem end to end, the work gets passed around, agent to agent, department to department. Each handoff is another chance to lose context, repeat a question, or hand off an answer that’s fast but wrong. The veteran ends up interacting with what should be a lifeline far more than they ever should.

Here is why it continues: We measure the activity, not the outcome. A customer satisfaction score can look excellent on a dashboard while the person behind it called four or more times, re-explained the same issue to four agents and two systems, and got the wrong answer twice. A call gets routed to the wrong place, labeled resolved and closed. Ten days later the veteran discovers the information was wrong and calls back, but because the measurement window was only three to seven days, the original failure was never recorded as a failure at all. The dashboard stays green. Multiply that across millions of interactions and you get a system that reports success while people quietly call back, again and again. Because we measure activity, we miss the goal of an actual resolution.

]]>

Now, agencies are moving artificial intelligence into these same pipelines, claims intake, contact centers, case triage. Done right, automation is exactly the fix: an AI that pulls the eligibility record my father’s case needed, that carries context across a handoff, that takes ownership a human queue couldn’t. Done the way we measure today, it does the opposite. It automates the bounce. It closes a fast wrong answer before anyone checks whether it was right, at a scale and speed no call center ever could. The same blind spot, now running at machine speed.

Federal agencies are making real, hard-won progress on speed. The VA, for instance, cut its claims backlog to roughly 70,000 after processing a record 3 million-plus claims in fiscal 2025. But speed is not the same as a correct answer, and most agencies still cannot show, end to end, that a given case was decided correctly.

Federal policy already gestures at the gap but stops short of closing it. The Office of Management and Budget’s Memorandum-25-22 directs agencies to monitor AI performance and safeguard taxpayer dollars when they buy these systems. The Government Accountability Office’s AI Accountability Framework names performance and monitoring among its core principles. OMB’s M-26-04 ties federal AI to public trust. Each requires “performance.” None defines it as the only thing that matters to a veteran: Did the problem actually get solved?

Agencies can close the gap in the contracts they are writing right now, while the standards are still being set. Two requirements would do most of the work.

First, measure resolution, not activity. Pay for first-contact resolution, for claims that don’t reopen, for determinations that hold up on appeal, and extend the measurement window past seven days, so a wrong answer that surfaces on day 10 counts against the system instead of vanishing from it. The layman’s version belongs in the contract itself: The AI must be an asset to the citizen and to the human resolving the issue, accurate and reliable enough for a true resolution, not just a faster call.

Second, require provenance. Every agentic system should log which model made or recommended which decision, on what data. Without that record, accountability is a guess; with it, a wrong determination can be audited, appealed and traced. That is what turns GAO’s “monitoring” principle and M-26-04’s “public trust” from aspiration into something an agency can enforce.

A friend of mine would put it more plainly than any of this: We’re shooting at the wrong hoop. We are dealing with veterans and citizens who deserve the highest possible resolution, and for them, speed was never the measure of service, accuracy is. Some of these cases are matters of life and death — most aren’t — all of them deserve a system built to solve the root of the problem rather than to get you off the phone.

]]>

My father’s answer was sitting in the database the whole time. No veteran, or anyone else, should need years and a lawyer to reach an answer the government already has.

Brandon Burdin is the founder of MarginSignal OS. He is the caregiver to his father, a disabled veteran.

Copyright
© 2026 Federal News Network. All rights reserved. This website is not intended for users located within the European Economic Area.