The Future of Trustworthy Agents Is Between Testing and Proof

The future of trustworthy enterprise AI sits between agents grading agents and formal verification.

The Future of Trustworthy Agents Is Between Testing and Proof

The enterprise AI industry has settled on a strange solution to unreliable AI: ask another AI whether the first AI can be trusted.

One agent produces an answer. Another critiques it. A third scores the result. The system runs through an evaluation suite, the dashboard turns green, and somebody declares it ready for production.

The process can improve an AI system. It cannot make the system's claims true.

That difference matters now that AI is moving beyond emails and document summaries. Companies want agents to recommend how they should allocate capital, hire people, set prices, schedule maintenance, and manage risk. A made-up fact in a summary wastes time. Put that same fact inside a budget or workforce plan and it can move money and people before anyone spots it.

We need a more useful definition of trust. Having agents grade other agents is not enough. Formally verifying everything a company does is neither practical nor, for many business questions, possible.

There is a workable space between the two. Test the behavior that cannot be specified completely. Prove the claims that admit proof. Trace the inputs that can change a decision. Make uncertainty visible. When the evidence runs out, the system should say so.

Agents checking agents

LLM judges, multi-agent critics, red teams, and end-to-end evaluations all have a place. They uncover failures, help compare versions, and give us evidence about how an agent is likely to behave.

They are still probabilistic evaluators.

The model grading an answer may share the assumptions and blind spots of the model that produced it. Research on LLM judges has found position bias and preferences for certain response styles. A second model may be more critical than the first, but that does not make it an independent source of truth.

Long agent workflows make this worse. An early reasoning error can become a premise for everything that follows. Other agents may elaborate it, write code around it, retrieve evidence that appears to support it, and eventually agree with one another. The original mistake now looks thoroughly researched.

More agents mean more scrutiny. They do not give us verification.

A passed evaluation tells us how a system performed on selected cases under selected conditions. It cannot prove that the next recommendation is correct, that every source is real, or that the model behind a decision preserves what a person asked for.

Testing gives us evidence. We should use it as evidence.

What formal verification can do

Formal verification offers something much stronger. Systems such as Lean can check mathematical proofs with extraordinary rigor, and formal methods already do useful work in software and infrastructure.

Amazon is a good example. Its automated reasoning teams apply formal logic to bounded questions in authorization, security, infrastructure, and policy. AWS IAM Access Analyzer can determine whether a policy grants unintended access. The question is precise, the consequences matter, and a machine-checkable answer is worth the effort.

That effort is substantial. Mature formal libraries such as Mathlib and the libraries used to reason about physics and software systems take years of specialized work. Formalization is not a switch we can flip after an agent returns an answer. People have to define the objects, encode the assumptions, state the properties, and build the library that makes the proof possible.

Business questions rarely arrive in that form. They sound more like this:

How should we grow in Houston?

Before anyone can prove anything, someone has to decide what "grow" means. They have to identify the choices the company can make, the limits it faces, the relevant time period, and the evidence for the expected effect of each action.

Much of the risk sits inside that translation.

A theorem prover checks a proof of the statement it receives. A solver finds an optimum for the program it receives. Neither one can tell us, on its own, whether that statement or program represents the decision the person meant to make.

You can have a valid proof of the wrong problem.

Trust needs an object

People talk about "a trustworthy model" or "a trustworthy agent" as if trust were a badge attached to a product. That is too vague for a consequential decision. We need to say which claim we trust, in which context, and why.

Take an AI-generated operating plan. Did its numbers come from real sources? Did it label estimates as estimates? Does its mathematical model preserve the stated objective and constraints? Is the plan feasible? Is the arithmetic correct? Under the declared assumptions, is it optimal? How far can an uncertain estimate move before the recommendation changes? Who approved the judgments that a machine could not check?

Those questions are often compressed into one trust score, perhaps produced by another language model. They should not be.

Each question calls for a different kind of assurance. Provenance records where a number came from. Deterministic code can enforce schemas and policy rules. Exact arithmetic can recheck feasibility and some optimality claims. Evaluations can probe open-ended behavior. Sensitivity analysis can identify the assumptions that could change a decision. A human owner still has to decide whether the model fits the situation.

The cost of a mistake should determine how much assurance we require. Nobody needs a proof before an AI suggests a meeting title. A staffing plan, credit decision, or maintenance schedule deserves a higher bar.

The NIST AI Risk Management Framework and its Generative AI Profile point in this direction. NIST treats trustworthiness as a matter of context, risk tolerance, provenance, measurement, human oversight, and safe failure. It also tells organizations to account for risks they cannot measure. That is far more useful than a generic claim that an agent passed its benchmark.

Let the system refuse

AI products are rewarded for answering. A blank field looks like a defect, while a polished recommendation looks like capability. For high-stakes decisions, this incentive is backwards.

I would rather deploy a system that answers 70 percent of consequential questions and clearly refuses the rest than one that answers 95 percent while hiding occasional errors. The second system will look better in a demo. It is also the one people will learn to trust right before it fails them.

A responsible agent has to distinguish an answer from a justified answer.

Refusal is a feature, not a temporary flaw that larger models will eliminate. Some decisions will always involve missing evidence, effects the available data cannot identify, contested objectives, or judgments that belong to a person. When a critical input is missing, the truthful response is to name it and stop.

Enterprises should still move quickly. The capabilities are real, and waiting forever is not a strategy. But no vendor can honestly promise that its AI will produce only positive outcomes. A better standard is measurable value while keeping unresolved risk within the organization's tolerance, especially when money, livelihoods, safety, or critical operations are involved.

The space between agents and Lean

A practical assurance system uses several methods together. It evaluates open-ended behavior empirically. It applies machine-checkable verification to bounded claims. It records the origin of consequential inputs. It leaves modeling judgment with an accountable human. If an important obligation remains open, it refuses to issue a verified answer.

At KNOWIDEA, this is the role we want VeriPIE to play. VeriPIE does not try to formally verify an entire business decision or declare an AI system trustworthy in every setting. It translates an operational question into an explicit program, records where the influential numbers came from, checks the claims that allow exact verification, and exposes the assumptions and judgments that remain. If the required obligations do not close, it withholds a verified recommendation.

VeriPIE is one implementation of a broader idea: an enterprise does not need to formally verify every AI output. It should demand the strongest assurance each consequential claim can reasonably support.

The questions to ask before deployment

Companies are rushing to catch the agentic AI wave. The pressure is understandable, but many are deploying first and working out what "trust" means afterward. That order is going to hurt them.

Before an agent enters a consequential workflow, ask what claims people will rely on. Find out what evidence supports each one. Separate the claims that were tested from those that were independently checked. Make sure every number that could change the decision has a traceable origin. Name the assumptions and the decisions that still require a human owner. Decide in advance when the system must refuse and who remains accountable when it is wrong.

The companies that benefit most from AI will not be the ones that trust it most. They will be the ones that know exactly where their trust ends.

Trustworthy AIEnterprise AIVerificationVeriPIE

See how KNOWIDEA works for your team.

Book a 30-minute demo and bring us the decision you are working through.