Aplomb 1: open-weights 5.3B decision model, 1M context, text/image/video/audio in one request, #1 among 4B models on the Decision Index
Summary
EmpirioLabs released Aplomb 1, a small open-weights model built to make decisions instead of writing prose. It answers yes or no, choice, score, and tool-selection questions, and it accepts text, JSON, images, video, and audio together in a single request with a very long context window.
Its distinguishing feature is that it returns a probability for each tool and for each enum and boolean argument, so an agent can act on confident calls and pass uncertain ones to a larger model. The publisher says it built the model on Qwen base models with an added decision head, and it ships the weights under its own license, which is free for research, evaluation, personal use, and internal use at smaller companies.
Why it matters
Decision models like this sit in the routing layer of AI agents, where a system chooses which tool to call, which document answers the question, and whether the input contains an answer at all. If cheap, fast models with calibrated probabilities take over that step, a brand's content gets accepted or dropped by a confidence threshold before any larger model reads it. Pages, product data, and media that state facts plainly and unambiguously are more likely to clear that threshold, and content that leaves the answer uncertain is more likely to be passed over or scored as not containing an answer.
Source
- docs.empiriolabs.ai/models/aplomb-1 the item itself, published by r/LocalLLaMA, created by empiriolabsai
- empiriolabs.ai/blog/introducing-aplomb-1 Reproducible first-party demonstration, first party
- huggingface.co/empiriolabsai/aplomb-1 Original dataset with disclosed methods
- docs.empiriolabs.ai/models/aplomb-1.md Named expert analysis, or a first-party claim about itself, first party
- reddit.com/r/LocalLLaMA/comments/1wz96qs/aplomb_1_openweights_53b_decision_model_1mwhere AIOTruth found it
What was checked
- canonical URL taken from the page itself
- read the item page (200)
- https://empiriolabs.ai/blog/introducing-aplomb-1: supports the item, tier 2 (Reproducible first-party demonstration)
- https://docs.empiriolabs.ai/models/aplomb-1: supports the item, tier 8 (Named expert analysis, or a first-party claim about itself)
- https://docs.empiriolabs.ai/models/aplomb-1.md: supports the item, tier 8 (Named expert analysis, or a first-party claim about itself)
- https://huggingface.co/empiriolabsai/aplomb-1: supports the item, tier 5 (Original dataset with disclosed methods)
- https://platform.empiriolabs.ai/dashboard/playground?model=aplomb-1: does not mention the item, tier 8 (Named expert analysis, or a first-party claim about itself)
Limitations
- Reddit engagement is a discovery signal here, not evidence that the claim is correct.
AIOTruth judgment
The Confirmed status reflects that the launch post, the documentation, and the Hugging Face model card were all fetched and all describe the same release, with weights and a public benchmark run available for anyone to inspect. Evidence and usefulness score high because the model can be downloaded and run, and originality scores high because per-argument probabilities across mixed media inputs are an unusual combination. Relevance is lower because this is agent infrastructure, not a direct discovery or citation surface. The benchmark standings are the publisher's own run of the kit, and the Reddit thread is a discovery signal, not evidence that the claims are correct.
How this score was calculated
| Dimension | Weight | Score | What produced it |
|---|---|---|---|
| AIO relevance | 30% | 6 | ai mode, rag; 4 scope question(s) matched |
| Usefulness | 30% | 9 | 3 artifact(s), 5 actionable marker(s), 1 measured figure(s) |
| Evidence | 25% | 9 | Reproducible first-party demonstration; 4 verified source(s) across 2 domain(s) |
| Originality | 15% | 9 | carries original testing or data; closest archive match 0 |
aioRelevance x 0.30 + usefulness x 0.30 + evidence x 0.25 + originality x 0.15. The rubric is published in full on the Editorial Method page. Scoring is deterministic: the same item scores the same on every run.
Review the documentation Open the source
Found something wrong here? AIOTruth corrects material errors openly. Have something we should review? Submit a find.