Original measurement · 26 August 2026·12 min read

We asked four AI engines to recommend AI tools. They agreed 12% of the time.

Three independent models converge on a presentation tool whose best-known domain now returns a 404. This is original measurement, not commentary: 100 captures across ten categories, taken on one day, with the protocol and the raw numbers published below.

Snapshot: 26 August 2026 Re-measure: November 2026
Bottom line up front

How often do engines agree?

Proportion of each category's tools that all three memory engines independently named. Same track length on every row, because what is being compared is the ratio.

AI presentation tools6 of 19 tools32%
AI transcription & meetings6 of 3119%
AI coding assistants5 of 2520%
AI recruiting software2 of 1811%
AI voice & text to speech3 of 2711%
AI writing tools3 of 2612%
AI chatbots & assistants3 of 2811%
AI short-form clipping2 of 268%
AI image generators1 of 333%
AI video generators1 of 353%

Tools named by all three memory engines, as a share of every tool any of them named in that category. 268 distinct tools across 10 categories, 90 captures, 26 August 2026.

The spread matters more than the average. In presentation software the engines behave like they are reading the same list. In image and video generation they behave like they are describing different industries: 33 and 35 tools named, exactly one of each shared by all three. A category with 3% agreement has no consensus to appeal to, which means the engine a buyer happens to open largely determines the shortlist they will ever see.

The tool all three agreed on has a dead front door

Presentation software was the highest-consensus category in the study. One of its six consensus tools no longer answers at the domain it was known by.

Tome · recommended 8 times · product closed 30 April 2025
GET https://tome.app
→ HTTP 404
→ "The deployment could not be found on Vercel.
   DEPLOYMENT_NOT_FOUND"
→ 104 characters of content

Tome was named by all three memory engines, across multiple phrasings of the question, making it the fourth most-recommended tool in the entire 268-tool set. It sat alongside Gamma, Beautiful.ai and Slidebean as an engine-consensus pick.

The grounded engine did not name it once. Asked the same question with web search enabled, it returned Gamma, Alai, Plus AI and Canva.

The product was discontinued sixteen months before this snapshot. Tome closed its presentation product on 30 April 2025, pivoted to sales automation, and the brand was sold on; our own July 2026 reporting recorded that you could no longer sign up. The engines are not naming a product that moved. They are naming one that was retired well over a year before they were asked.

One caveat we will not paper over. A live site at tomeapp.ai sells an AI presentation tool under the Tome name and returned 200 with 8,217 characters when we checked. We cannot establish from a website whether it is connected to the original company, and we assert nothing either way. It does not change the measurement: engines emit a brand name, not a URL, and the domain that brand was known by answers with a Vercel error.

This is the clearest illustration we found of a simple point: agreement between models is not evidence about the world. Three systems trained on overlapping corpora will overlap in their stale entries too, and the more confidently they converge, the more convincing the answer looks. Consensus measures shared training, not current reality. It is also a reminder of what these systems actually emit: a name, unaccompanied by any check that the name still resolves to something a buyer can reach.

A second case, and it is subtler

The same pattern appears where a product did not die but changed its name.

Memory engines

coding assistants
CodeWhisperer
named by all 3
5 mentions
The Amazon brand that no longer ships under that name.
A buyer searching for it finds a redirect, at best.

Grounded engine

same question, web search on
Amazon Q Developer
named once
current product
Returned alongside Claude Code, ChatGPT and Codex.
Grounding produced the name a buyer can actually act on.

Two independent errors, one dead site and one retired brand, both present in model memory and both absent from grounded search. That is a small sample and we are not claiming grounding fixes everything. It is enough to say that on the two verifiable staleness errors this study found, search corrected both.

Grounded and memory answers barely overlap

Of 43 tools the grounded engine recommended across ten categories, fewer than half appear anywhere in the memory engines' far larger set.

CategoryGrounded namedMemory namedIn bothOverlap
AI short-form clipping4264100%
AI presentation tools419375%
AI voice & text to speech327267%
AI writing tools326267%
AI transcription & meetings731457%
AI image generators433250%
AI recruiting software318133%
AI coding assistants625117%
AI video generators53500%
AI chatbots & assistants42800%
Two categories share nothing at all. For AI video generators, grounded search returned Veo, Kling, DeeVid, Adobe Firefly and Luma; the memory engines named 35 tools and not one of those five. The same holds for chatbots. A vendor can be the memory consensus and invisible to search, or the reverse, and neither position tells you anything about the other.

Note the direction of the surprise. The memory engines cast a far wider net, 268 tools to 43, so you would expect the smaller grounded set to be almost entirely contained within it. It is not. Grounded search is not returning a subset of what the models remember; it is frequently returning a different set of names.

How many recommended tools are actually alive?

We resolved every tool with an unambiguous domain and requested its homepage.

OutcomeCountWhat it means
Alive36200 response with real content
Domain dead1tome.app: 404, Vercel error page, 104 characters. A product under that name is live at another domain
Unreadable8Bot protection or JavaScript-only. Nothing measured, never counted as dead
Untestednot countedTools with no unambiguous domain were excluded rather than guessed
Our first pass reported five dead products. Four were wrong, and the correction is instructive. The detector matched the phrase "coming soon" anywhere on the page, which appears in ordinary marketing copy on healthy sites: Pitch, Writesonic, Runway and Rask AI all returned 200 with between four and fifteen thousand characters of live product content. A dead site is not a phrase, it is a phrase on an empty page. The test now requires a platform error string and a near-empty response. We are reporting this because a study about machines confidently repeating stale claims has no business making four of its own.
Correction · 26 August 2026

This page was first published stating that the consensus tool does not exist. That was too strong as written. The product was discontinued on 30 April 2025 and tome.app returns a Vercel 404, but a live site at tomeapp.ai sells an AI presentation tool under the same name, and we cannot establish from a website whether the two are connected.

The claim has been narrowed to what we actually measured: the domain most associated with the brand the engines named answers with an error. We are recording the change rather than quietly editing it, because a page arguing that machines repeat unchecked claims has to hold itself to the same standard.

Method, and what this does not show

Protocol · 26 August 2026

This uses the published protocol of The AI Recommendation Audit 2026 (DOI 10.5281/zenodo.20767878, CC-BY), which measured 16 B2B categories, pointed at ten AI-tool categories instead. Same collector, same prompts, same parser; only the category bank changed.

Engines. Three model-memory engines answered three buying questions per category: GPT-OSS 120B, Qwen 3.6 and Cohere Command-A (90 captures). One grounded engine, Gemini 2.5 Flash with Google Search, answered one question per category (10 captures).

Limits, stated plainly

The grounded sample is one engine and one query per category. A daily grounding quota made a full three-query pass impossible on the day, and a second grounded engine hit its rate limit at zero captures. The grounded column is therefore a probe, not a survey, and every grounded claim here should be read as such.

Memory had three times the query surface. That explains why memory named more tools; it does not explain why the smaller grounded set falls outside it, which is the finding.

This is a dated snapshot, not a ranking. AI answers are volatile. Re-running the published protocol on another day will produce different sets, which is itself part of the point.

We did not test answer quality. Nothing here says the recommended tools are good, only which names come back and whether those names still resolve to a product.

What to do with this

If you are choosing a tool
Ask twice
two engines, minimum

One engine's answer is a sample, not a verdict. At 12% agreement, a second engine will usually show you a mostly different shortlist.

If a recommendation matters
Open the site
ten seconds

The most-agreed-upon pick in the highest-consensus category was a 404. Confidence and convergence are not evidence of existence.

If you sell a tool
Measure per engine
not once, repeatedly

Being the memory consensus and being findable by search are separate properties. In two categories they had zero overlap.

Frequently asked questions

Do different AI engines recommend the same tools?
Mostly not. Across 10 categories on 26 August 2026, three memory engines named 268 distinct tools and only 32 (11.9%) were named by all three. Agreement ranges from 32% in AI presentation software down to 3% in image and video generation.
Do AI engines recommend products you cannot reach?
Yes, at least once here. All three memory engines recommended Tome for presentations, 8 mentions, the fourth most-recommended tool in the study, and tome.app returns HTTP 404 with a Vercel DEPLOYMENT_NOT_FOUND error. A product called Tome App AI is live at tomeapp.ai; we cannot establish from a website whether it is the same company and do not assert it. Engines emit brand names, not URLs, so a recommendation can still lead to a dead address. Of 45 tools tested, 36 were reachable, 8 unreadable behind bot protection, 1 with a dead canonical domain.
Does web search fix stale AI recommendations?
In both verifiable cases here it did. The grounded engine never named Tome, and it returned Amazon Q Developer for coding assistants where all three memory engines returned the retired Amazon CodeWhisperer brand. That is two corrections out of two staleness errors found, on a small grounded sample.
How much do grounded and memory AI answers overlap?
Less than half. The grounded engine named 43 tools and only 19 (44.2%) appeared anywhere in the memory engines' 268. In AI video generators and AI chatbots the two shared no recommendation at all.
Why does it matter which AI engine a buyer asks?
Because the answer sets barely overlap, the engine a buyer opens largely determines the shortlist they see. A vendor can be the consensus pick of three models and invisible to a fourth. For buyers, one engine's answer is a sample. For vendors, AI visibility must be measured per engine rather than assumed.
How was this measured?
Using the published protocol of The AI Recommendation Audit 2026 (DOI 10.5281/zenodo.20767878) pointed at 10 AI-tool categories. Three memory engines answered three questions per category (90 captures); one grounded engine answered one per category (10 captures). Snapshot 26 August 2026.

Related reading on Nesyona

Save
Dashboard

From our network

Best AI Tools for Amazon Sellers - bagengine.comBest AI Courses 2026 - edubracket.comBest Accounting Software for Online Sellers - ceocult.com