2 months (meme-d) research, Do People Notice the Difference Between AI Models?
Summary
A r/LocalLLaMA poster ran an informal two month experiment to see whether everyday users can tell high end AI models apart. Thirteen people were told they were getting free, time limited access to the latest ChatGPT model through the poster's own website. Eight of them appear to have used it. Behind a gray themed OpenWebUI interface with the admin panels hidden, the answers actually came from locally hosted open models, including a Qwen model.
The poster started with a mix of local models and OpenRouter, then dropped OpenRouter after the first two weeks because a single consumer graphics card handled the light, bursty usage. The stated goal was to see whether participants would notice meaningful differences between models, complain about quality, or form preferences without knowing which model was answering. The author says up front that the sample is too small to be more than curiosity, and that participants gave permission to publish afterward.
Why it matters
If ordinary users cannot tell which model is answering, then the brand name on an AI interface is carrying much of the trust, not the model underneath. For anyone trying to get found, understood, cited, or recommended by AI, that means the model behind a given assistant can change without users noticing, and how a brand is described can change with it. Checking how a brand is represented needs to cover several models, including open ones served through third party interfaces, not only the best known assistant.
Source
- huggingface.co/Lorbus/Qwen3.6-27B-int4-AutoRound the item itself, published by r/LocalLLaMA, created by Altruistic_Heat_9531
- github.com/eugr/spark-vllm-docker Public source code or repository
- huggingface.co/docs/inference-providers/index Original dataset with disclosed methods, first party
- reddit.com/r/LocalLLaMA/comments/1x0etwe/2_months_memed_research_do_people_notice_thewhere AIOTruth found it
What was checked
- canonical URL taken from the page itself
- read the item page (200)
- https://github.com/intel/auto-round: does not mention the item, tier 6 (Public source code or repository)
- https://github.com/eugr/spark-vllm-docker: supports the item, tier 6 (Public source code or repository)
- https://github.com/vllm-project/vllm: does not mention the item, tier 6 (Public source code or repository)
- https://huggingface.co/docs: does not mention the item, tier 5 (Original dataset with disclosed methods)
- https://huggingface.co/docs/safetensors/index: does not mention the item, tier 5 (Original dataset with disclosed methods)
- https://huggingface.co/docs/inference-providers/index: supports the item, tier 5 (Original dataset with disclosed methods)
Limitations
- Reddit engagement is a discovery signal here, not evidence that the claim is correct.
AIOTruth judgment
The item is rated Supported, with evidence and originality as its strongest dimensions. It is a first hand account that discloses its setup, its hardware, its sample size, and its own weaknesses, and a blind substitution test on real users is uncommon. Usefulness is solid because the design is simple enough for others to repeat. Relevance is lower because the experiment is about model perception in general, not brand discovery specifically. The recorded limitation applies: Reddit engagement is a discovery signal here, not evidence that the claim is correct, and the author describes the work as informal.
How this score was calculated
| Dimension | Weight | Score | What produced it |
|---|---|---|---|
| AIO relevance | 30% | 7 | ai mode; 4 scope question(s) matched |
| Usefulness | 30% | 8 | 3 artifact(s), 6 actionable marker(s), 0 measured figure(s) |
| Evidence | 25% | 9 | Original dataset with disclosed methods; 2 verified source(s) across 2 domain(s) |
| Originality | 15% | 9 | carries original testing or data; closest archive match 0 |
aioRelevance x 0.30 + usefulness x 0.30 + evidence x 0.25 + originality x 0.15. The rubric is published in full on the Editorial Method page. Scoring is deterministic: the same item scores the same on every run.
Read the source Open the source
Found something wrong here? AIOTruth corrects material errors openly. Have something we should review? Submit a find.