We asked four AI engines to recommend AI tools. They agreed 12% of the time.
Three independent models converge on a presentation tool whose best-known domain now returns a 404. This is original measurement, not commentary: 100 captures across ten categories, taken on one day, with the protocol and the raw numbers published below.
- Cross-engine agreement is 11.9%. Three memory engines named 268 distinct tools; only 32 were named by all three.
- All three recommended Tome, and tome.app returns HTTP 404. A product under that name operates at a different domain, so the recommendation resolves to a dead front door rather than to nothing.
- Grounded search dropped it. The web-searching engine named Tome nowhere, and returned the current Amazon Q Developer where memory returned the retired Amazon CodeWhisperer.
- Grounded and memory answers overlap 44.2%. In AI video generators and AI chatbots they shared no recommendation at all.
- Agreement is wildly category-dependent: 32% in presentation tools, 3% in image and video generation.
How often do engines agree?
Proportion of each category's tools that all three memory engines independently named. Same track length on every row, because what is being compared is the ratio.
Tools named by all three memory engines, as a share of every tool any of them named in that category. 268 distinct tools across 10 categories, 90 captures, 26 August 2026.
The tool all three agreed on has a dead front door
Presentation software was the highest-consensus category in the study. One of its six consensus tools no longer answers at the domain it was known by.
→ HTTP 404
→ "The deployment could not be found on Vercel.
DEPLOYMENT_NOT_FOUND"
→ 104 characters of content
Tome was named by all three memory engines, across multiple phrasings of the question, making it the fourth most-recommended tool in the entire 268-tool set. It sat alongside Gamma, Beautiful.ai and Slidebean as an engine-consensus pick.
The grounded engine did not name it once. Asked the same question with web search enabled, it returned Gamma, Alai, Plus AI and Canva.
The product was discontinued sixteen months before this snapshot. Tome closed its presentation product on 30 April 2025, pivoted to sales automation, and the brand was sold on; our own July 2026 reporting recorded that you could no longer sign up. The engines are not naming a product that moved. They are naming one that was retired well over a year before they were asked.
One caveat we will not paper over. A live site at tomeapp.ai sells an AI presentation tool under the Tome name and returned 200 with 8,217 characters when we checked. We cannot establish from a website whether it is connected to the original company, and we assert nothing either way. It does not change the measurement: engines emit a brand name, not a URL, and the domain that brand was known by answers with a Vercel error.
This is the clearest illustration we found of a simple point: agreement between models is not evidence about the world. Three systems trained on overlapping corpora will overlap in their stale entries too, and the more confidently they converge, the more convincing the answer looks. Consensus measures shared training, not current reality. It is also a reminder of what these systems actually emit: a name, unaccompanied by any check that the name still resolves to something a buyer can reach.
A second case, and it is subtler
The same pattern appears where a product did not die but changed its name.
Memory engines
5 mentions
Grounded engine
current product
Two independent errors, one dead site and one retired brand, both present in model memory and both absent from grounded search. That is a small sample and we are not claiming grounding fixes everything. It is enough to say that on the two verifiable staleness errors this study found, search corrected both.
Grounded and memory answers barely overlap
Of 43 tools the grounded engine recommended across ten categories, fewer than half appear anywhere in the memory engines' far larger set.
| Category | Grounded named | Memory named | In both | Overlap |
|---|---|---|---|---|
| AI short-form clipping | 4 | 26 | 4 | 100% |
| AI presentation tools | 4 | 19 | 3 | 75% |
| AI voice & text to speech | 3 | 27 | 2 | 67% |
| AI writing tools | 3 | 26 | 2 | 67% |
| AI transcription & meetings | 7 | 31 | 4 | 57% |
| AI image generators | 4 | 33 | 2 | 50% |
| AI recruiting software | 3 | 18 | 1 | 33% |
| AI coding assistants | 6 | 25 | 1 | 17% |
| AI video generators | 5 | 35 | 0 | 0% |
| AI chatbots & assistants | 4 | 28 | 0 | 0% |
Note the direction of the surprise. The memory engines cast a far wider net, 268 tools to 43, so you would expect the smaller grounded set to be almost entirely contained within it. It is not. Grounded search is not returning a subset of what the models remember; it is frequently returning a different set of names.
How many recommended tools are actually alive?
We resolved every tool with an unambiguous domain and requested its homepage.
| Outcome | Count | What it means |
|---|---|---|
| Alive | 36 | 200 response with real content |
| Domain dead | 1 | tome.app: 404, Vercel error page, 104 characters. A product under that name is live at another domain |
| Unreadable | 8 | Bot protection or JavaScript-only. Nothing measured, never counted as dead |
| Untested | not counted | Tools with no unambiguous domain were excluded rather than guessed |
This page was first published stating that the consensus tool does not exist. That was too strong as written. The product was discontinued on 30 April 2025 and tome.app returns a Vercel 404, but a live site at tomeapp.ai sells an AI presentation tool under the same name, and we cannot establish from a website whether the two are connected.
The claim has been narrowed to what we actually measured: the domain most associated with the brand the engines named answers with an error. We are recording the change rather than quietly editing it, because a page arguing that machines repeat unchecked claims has to hold itself to the same standard.
Method, and what this does not show
This uses the published protocol of The AI Recommendation Audit 2026 (DOI 10.5281/zenodo.20767878, CC-BY), which measured 16 B2B categories, pointed at ten AI-tool categories instead. Same collector, same prompts, same parser; only the category bank changed.
Engines. Three model-memory engines answered three buying questions per category: GPT-OSS 120B, Qwen 3.6 and Cohere Command-A (90 captures). One grounded engine, Gemini 2.5 Flash with Google Search, answered one question per category (10 captures).
The grounded sample is one engine and one query per category. A daily grounding quota made a full three-query pass impossible on the day, and a second grounded engine hit its rate limit at zero captures. The grounded column is therefore a probe, not a survey, and every grounded claim here should be read as such.
Memory had three times the query surface. That explains why memory named more tools; it does not explain why the smaller grounded set falls outside it, which is the finding.
This is a dated snapshot, not a ranking. AI answers are volatile. Re-running the published protocol on another day will produce different sets, which is itself part of the point.
We did not test answer quality. Nothing here says the recommended tools are good, only which names come back and whether those names still resolve to a product.
What to do with this
One engine's answer is a sample, not a verdict. At 12% agreement, a second engine will usually show you a mostly different shortlist.
The most-agreed-upon pick in the highest-consensus category was a 404. Confidence and convergence are not evidence of existence.
Being the memory consensus and being findable by search are separate properties. In two categories they had zero overlap.