NesyonaResearch  ›  The AI Recommendation Audit 2026
Original research · 2026-08-26

4 AI engines, one question, little agreement on the best AI tool

We asked the AI assistants your buyers actually use which software they recommend, across 10 go-to-market categories. Here is what they say, who is invisible, and why the answer depends on which AI you ask.

20%of 10 categories where all 4 engines named the same #1 tool
4
AI engines
10
AI-tool categories
30
buying queries
550
recommendations
256
tools named
Leaderboard

The AI-Visibility leaderboard

Every tool's position-weighted recommendation frequency across all engines and categories, normalized 0 to 100. A #1 pick counts more than a #5; breadth across categories is rewarded.

ToolAI-VisibilityScoreMentions#1 pickCats
1Otter.ai
10011102
2Jasper
89.91092
3GitHub Copilot
78.7971
4Beautiful.ai
65.6951
5Descript
63.11614
6Gamma
56.6941
7Opus Clip
51.6651
8ElevenLabs
49.2551
9ChatGPT
48.8743
10HireVue
45.9731
11Pictory
42.6632
12Cursor
40.6721
13Tome
36.9811
14Midjourney
35.2531
15Synthesia
29.7622
16Copy.ai
29.5701
17Claude
26.2512
18Copilot
25.4612
19Runway
24.6321
20Gemini
23502
21Stable Diffusion
21.6601
22Tabnine
20701
23Rev
19.7312
24Greenhouse
19.7221
25Resemble AI
18311

Top 25 of 256 measured tools. Linked tools have a full AI-Visibility report card. Complete ranking in the open dataset.

The centerpiece

Do different AI engines agree on who to recommend?

Mostly, no. Below is each engine's top pick in every category. The five engines named the same single best tool in only 20% of categories, disagreeing in 8 of 10. Which AI a buyer happens to ask changes which tool they're told to buy.

web-grounded engine (reads the live web)model-memory engine (recalls training data) engines disagree on #1
CategoryGPT-OSSmemoryQwen 3.6memoryCohere Command-AmemoryGemini 2.5 FlashgroundedAgree?
AI image generatorsMidjourneyMidjourneyMidJourney v7GPT Image 2
AI presentation toolsBeautiful.aiGammaBeautiful.aiGamma
AI video generatorsSynthesiaRunwaySynthesiaVeo
AI short-form video clippingOpus ClipOpus ClipDescriptOpus Clip
AI voice and text to speechElevenLabs PrimeElevenLabsResemble.aiElevenLabs
AI writing toolsJasperJasperJasperJasper=
AI coding assistantsGitHub CopilotGitHub CopilotGitHub CopilotClaude Code
AI transcription and meeting notesOtter.aiOtter.aiOtter.aiOtter.ai=
AI chatbots and assistantsChatGPTChatGPTOpenAI ChatGPT-5Rasa
AI recruiting softwareHireVueEightfold AIHireVuehireEZ
Why they diverge

Grounded engines and memory engines recommend different tools

The split above is not random. It tracks how each engine knows what it knows.

Model-memory engines

GPT-OSS and Qwen 3.6 and Cohere Command-A answer from training data. They reliably name the established, widely-written-about brands (the HubSpots and Mailchimps) and tend to miss anything newer than their training cut-off.

Web-grounded engines

Gemini 2.5 Flash read live search results, so they surface the current and trending winners (newer category leaders in email, sales, and AI-visibility tooling) that the memory engines never mention.

AI assistants increasingly answer "what's the best tool for X" before a buyer ever scans a page of links, which makes being named by AI a surface distinct from your Google ranking. The engines we queried are public and reproducible (GPT-OSS, Qwen 3.6, Cohere Command-A, Gemini 2.5 Flash), each asked the identical question set, with the full raw output in the open dataset below.

Source bias

Where AI gets its recommendations

When the grounded engines backed their answers with live web search, these are the domains they cited most. This is the map of where to be cited if you want AI to recommend you.

Is your tool in the short list?

Run the AI-Visibility check to see where your product lands when buyers ask AI for a recommendation.

Get your tool measuredDownload the data
FAQ

Frequently asked questions

Is AI visibility the same as SEO?
No. AI visibility is whether AI assistants name your tool when a buyer asks for a recommendation; SEO is your rank in the blue links. The two overlap only partly, so AI visibility is a separate, winnable race.
How is the AI-Visibility Score calculated?
It is a position-weighted count of how often a tool is recommended across the query set (a #1 recommendation counts more than a #5), normalized 0 to 100 and rewarded for breadth across categories. The full method and open data are on this page.
Does this rank product quality?
No. It measures what AI engines OUTPUT when asked, not which tool is objectively best. AI outputs reflect what their training data or the live web says, and can be wrong or biased. Treat it as a visibility snapshot, not an endorsement.
Why do the engines disagree so much?
Model-memory engines recall the established names from training data, while web-grounded engines read current pages and surface newer or trending tools. Same question, different information source, different answer.
How often is it updated?
It is a dated snapshot. AI answers shift, so the protocol is built to be re-run, and each refresh produces a delta showing how the recommendations moved.
Method & honesty floor

How this was measured

SOURCED: every recommendation is a recorded output of a live AI engine. Engines audited: GPT-OSS 120B (OpenAI open weights, model memory) [model-memory, 30 queries]; Qwen 3.6 (Alibaba, model memory) [model-memory, 30 queries]; Cohere Command-A (model memory) [model-memory, 30 queries]; Gemini 2.5 Flash (grounded) [web-grounded, 10 queries]. Across 10 AI-tool categories, the two model-memory engines answered all three phrasings per category; the three web-grounded engines (Gemini via API, ChatGPT and Perplexity via browser) answered one query per category, capturing live web sources where available. DERIVED: the AI-Visibility Score (position-weighted recommendation frequency, normalized 0-100), per-category leaderboards, cross-engine agreement, and source-citation tallies. HONESTY FLOOR: this measures what AI engines OUTPUT, not product quality or our opinion. Search-grounded engines reflect the live web; model-memory engines reflect training data and can name tools that do not exist or omit real leaders. Because the grounded engine is sampled less densely than the memory engines, the blended leaderboard leans slightly toward model-memory; the per-engine comparison below separates them. Outputs vary by phrasing, date, and engine; this is a dated snapshot, not a ranking endorsement. Browser passes (ChatGPT, Perplexity) are layered in as documented.

Engines: 4 Categories: 10 Queries: 30 Recommendations: 550 License: CC-BY 4.0 Snapshot: 2026-08-26

Open data: ai-recommendation-audit-2026.json, free to reuse with attribution to Nesyona. This measures what AI engines OUTPUT, not product quality; a dated snapshot, re-run to reproduce.

Cite this dataset

Couey, V. W. (2026). The AI Tools Recommendation Audit 2026 [Data set]. Nesyona.

Save
Dashboard