Content & Methods

Maat: Independent Deterministic Contract-Based Governance for Multi-Agent LLM Workflows

Supported Published September 27, 2026 Reviewed by AIOTruth September 29, 2026
7.5 AIOTruth Research Value Score

Summary

Maat is a runtime layer for multi-agent LLM workflows that checks every handoff between agents against a versioned workflow contract. No language model sits in the validation or scoring path, so the same handoff always gets the same verdict. The authors tested it across six controlled domain workflows with injected data-level defects and a deterministic rubric.

The paper is also a correction of its own first version. A post-publication audit found that three benchmark scorers counted any early halt as a prevented defect. After restricting the comparison to halts tied to a verified defect, rubric scores rose in five workflows and stayed flat in software development, and model-call cost fell where those halts came early. A hand review of every governed halt found that more than a third were false alarms caused by validator defects, and counting those as failed work puts the governed arm below the ungoverned arm in four of six workflows.

Why it matters

AI answers about brands increasingly come out of chained agents: one retrieves, one summarizes, one writes the recommendation. An error about a company at the first step can be accepted as context by every step after it. This paper shows that a deterministic contract check can stop the kind of defect a contract can express, such as a missing or malformed field, and that it does not detect hallucination. For anyone trying to be represented accurately by AI, the practical consequence is that structured, checkable facts are the part of a brand's information a pipeline can validate, while a wrong but well-formed claim still passes through.

Source

What was checked

  • canonical URL taken from the page itself
  • read the item page (200)
  • https://arxiv.org/abs/2609.34017: supports the item, tier 5 (Original dataset with disclosed methods)
  • https://arxiv.org/abs/2609.34017v1: supports the item, tier 5 (Original dataset with disclosed methods)
  • https://github.com/Lorelys/maat-benchmarks: supports the item, tier 6 (Public source code or repository)

Limitations

  • None recorded beyond what is stated above.

AIOTruth judgment

The Supported status and the evidence score of 9 rest on a first-party paper with disclosed methods and a public benchmark repository, and on the authors publishing an audit that weakens their own earlier result. Usefulness scores 9 because the findings on false alarms and halt attribution apply directly to anyone building or assessing agent pipelines. Originality at 7 reflects a deterministic alternative to LLM-based judges, tested in controlled workflows only. Relevance at 5 holds the overall Value Score to 7.5: the work concerns workflow reliability, and its link to how brands are found and cited is indirect. The authors state that the results do not establish universal correctness, hallucination detection, or model-independent effectiveness.

How this score was calculated

DimensionWeightScoreWhat produced it
AIO relevance30%5no beat terms matched; 3 scope question(s) matched
Usefulness30%92 artifact(s), 2 actionable marker(s), 1 measured figure(s)
Evidence25%9Original dataset with disclosed methods; 3 verified source(s) across 2 domain(s)
Originality15%7no original testing found; closest archive match 0

aioRelevance x 0.30 + usefulness x 0.30 + evidence x 0.25 + originality x 0.15. The rubric is published in full on the Editorial Method page. Scoring is deterministic: the same item scores the same on every run.

Read the source Open the source

Found something wrong here? AIOTruth corrects material errors openly. Have something we should review? Submit a find.

Run a check on AIOInsights