Maat: Independent Deterministic Contract-Based Governance for Multi-Agent LLM Workflows
Summary
Maat is a runtime layer for multi-agent LLM workflows that checks every handoff between agents against a versioned workflow contract. No language model sits in the validation or scoring path, so the same handoff always gets the same verdict. The authors tested it across six controlled domain workflows with injected data-level defects and a deterministic rubric.
The paper is also a correction of its own first version. A post-publication audit found that three benchmark scorers counted any early halt as a prevented defect. After restricting the comparison to halts tied to a verified defect, rubric scores rose in five workflows and stayed flat in software development, and model-call cost fell where those halts came early. A hand review of every governed halt found that more than a third were false alarms caused by validator defects, and counting those as failed work puts the governed arm below the ungoverned arm in four of six workflows.
Why it matters
AI answers about brands increasingly come out of chained agents: one retrieves, one summarizes, one writes the recommendation. An error about a company at the first step can be accepted as context by every step after it. This paper shows that a deterministic contract check can stop the kind of defect a contract can express, such as a missing or malformed field, and that it does not detect hallucination. For anyone trying to be represented accurately by AI, the practical consequence is that structured, checkable facts are the part of a brand's information a pipeline can validate, while a wrong but well-formed claim still passes through.
Source
- arxiv.org/abs/2609.34017 the item itself, published by arXiv, created by Uliana Elina
- arxiv.org/abs/2609.34017v1 Original dataset with disclosed methods, first party
- github.com/Lorelys/maat-benchmarks Public source code or repository
- arxiv.org/abs/2609.34017v1where AIOTruth found it
What was checked
- canonical URL taken from the page itself
- read the item page (200)
- https://arxiv.org/abs/2609.34017: supports the item, tier 5 (Original dataset with disclosed methods)
- https://arxiv.org/abs/2609.34017v1: supports the item, tier 5 (Original dataset with disclosed methods)
- https://github.com/Lorelys/maat-benchmarks: supports the item, tier 6 (Public source code or repository)
Limitations
- None recorded beyond what is stated above.
AIOTruth judgment
The Supported status and the evidence score of 9 rest on a first-party paper with disclosed methods and a public benchmark repository, and on the authors publishing an audit that weakens their own earlier result. Usefulness scores 9 because the findings on false alarms and halt attribution apply directly to anyone building or assessing agent pipelines. Originality at 7 reflects a deterministic alternative to LLM-based judges, tested in controlled workflows only. Relevance at 5 holds the overall Value Score to 7.5: the work concerns workflow reliability, and its link to how brands are found and cited is indirect. The authors state that the results do not establish universal correctness, hallucination detection, or model-independent effectiveness.
How this score was calculated
| Dimension | Weight | Score | What produced it |
|---|---|---|---|
| AIO relevance | 30% | 5 | no beat terms matched; 3 scope question(s) matched |
| Usefulness | 30% | 9 | 2 artifact(s), 2 actionable marker(s), 1 measured figure(s) |
| Evidence | 25% | 9 | Original dataset with disclosed methods; 3 verified source(s) across 2 domain(s) |
| Originality | 15% | 7 | no original testing found; closest archive match 0 |
aioRelevance x 0.30 + usefulness x 0.30 + evidence x 0.25 + originality x 0.15. The rubric is published in full on the Editorial Method page. Scoring is deterministic: the same item scores the same on every run.
Read the source Open the source
Found something wrong here? AIOTruth corrects material errors openly. Have something we should review? Submit a find.