Authentica
All resources

ProductSeptember 2026

Building an Invoice Audit Agent From an Empty Workspace

An unedited twenty-five-minute build. A freight invoice audit and dispute workflow, described in plain language and assembled into a governed operating model, deployed agents, and a first run that escalates to a human for the right reason.

Watch on YouTube

Most agent demos show you the finished thing. This one starts from an empty workspace and builds a freight invoice audit and dispute workflow while you watch, including the part where the first run stops and asks a human for a decision.

The premise is a shipper running a few thousand carrier invoices a month across about forty carriers, auditing roughly a tenth of them by hand. The ask, typed into the console in a couple of sentences, is for the audit to run as a system: coded against the contract, variances caught, disputes drafted, and anything the system is not sure about routed to a person. Nothing else exists at the start. No records, no agents, no workflows.

From a sentence to a build brief

The first conversation is with Kate, the assistant most people talk to first. She asks how invoices arrive today, then does the thing worth noticing: she declines to build, and hands off to the AI forward deployed engineer, which orchestrates builds. The FDE asks the questions an experienced implementer would ask. What does an auditor actually check? When is a variance worth disputing? Where does the “not sure” line sit?

Then it asks for the data, and gets it: a handful of invoice PDFs, an EDI document, an existing SOP, a contract rate table, an accessorial schedule, shipment records, and a fuel surcharge index. It reads all of it and maps it onto the operating model, the typed description of the entities, actions, tasks, and authority that the agents will run against.

What comes out is a build brief that states the scope before anything is built. Parse each invoice into typed charge lines. Check for duplicates. Look up the contract by carrier and lane. Audit linehaul and fuel surcharge. Write one carrier invoice record per invoice with its result: clean, variance, or unresolved. Roll the clean ones into a single payment release recommendation for the batch, draft a dispute for each variance, and escalate anything unresolved.

The brief also states, up front, which steps a human has to approve: releasing payment, sending each drafted dispute to a carrier, and deciding on unresolved exceptions. That list is part of the design, not a setting someone remembers to switch on later.

The change is proposed, reviewed, and versioned

The FDE hands the approved brief to the ontology author, which takes about five minutes to produce a proposed change: new entities, new actions, the tasks that compose the workflow, and the agents to run it. It arrives as a pending change in Studio, not as something already live.

That distinction is the whole point of the segment. The change sits there and can be argued with. In the video I trim the workflow down, check that the audit worker is not going to stop for human confirmation at every step, and confirm its trust level before accepting anything. Only then does it get merged and applied, through a git-style version control system with a compiler that validates the ontology against its own rules before it can reach the database.

Deploying the agents is a separate, deliberate step after that. Every customer gets an isolated agentic environment, so a worker and a supervisor are deployed into it explicitly. The supervisor is adversarial on purpose: models asked to grade their own work tend to think it went well.

The first run goes wrong in a useful way

Six invoices the system has never seen go in, along with an email and some shipment records. Kate builds a command plan, delegates it to the worker and the supervisor, and the workflow runs, forking after the audit step into the downstream paths: release payment, sign off the drafted disputes, or resolve an escalated exception.

All six come back unresolved, and the reason is instructive. The contracts were used to design the workflow but never loaded as data, so there is no contract on file to audit against. The agent does not guess a rate. It documents a plain-language reason per invoice and puts the decision in front of a person, which is exactly what the build brief said it would do.

One other thing surfaces mid-run. An action can legitimately be called many times inside a task, saving six invoice records in one step, for instance, and I had not reviewed how these particular tasks were designed. Where calling an action more than once would be wrong, the ontology can constrain it to exactly once. The guard is declarative, sits in the operating model, and is enforced at run time rather than requested in a prompt.

Why this is the demo we wanted to record

The objection we hear from IT is always the same, and it is a fair one: language models are nondeterministic, and they fail in unfamiliar ways. The answer is not a better prompt. It is a workflow whose scope, approval gates, and action constraints are written down, versioned, compiled, and enforced outside the model, so that the interesting question becomes what the system is permitted to do rather than what it happened to decide.

Twenty-five minutes, from an empty workspace to governed agents running a real workflow and escalating correctly. If you want to see the same thing on your own documents, book a demo.