Article's Content
Recommended, Not Just Mentioned: What We Learned Running One Buying Prompt 11 Times
There’s a number missing from every AI-visibility report your team has ever reviewed.
The headline metric in most AI-visibility reports, share of voice, only measures presence: how often your brand gets mentioned. It doesn’t capture whether the answer recommended you or pointed the buyer somewhere else.
But presence isn’t preference, and preference is what buyers act on. A brand can score well on presence and still lose the recommendation in the same answer.
We wanted to see how wide that gap actually gets.
So we had ten members of our marketing team to put the same question to ChatGPT or Claude:
“I’m on the hunt for a revenue intelligence platform that can help us replace our current stack. We’re onboarding 5 new SDRs next month, and we need them to have a tool that can help them hit their targets.”
Gong appeared in all 11 threads. But it was the clear recommendation in only three.
This piece breaks down why the recommendation moved, and how your team can start measuring it.
Early Preference Is Hard to Dislodge
The first recommendation an AI tool gives a buyer tends to stick.
According to 6sense’s 2025 Buyer Experience Report, 95% of buyers purchase from a vendor on their initial shortlist, and the vendor they contact first wins roughly 80% of the time. If your visibility dashboard already looks healthy, this is exactly the moment it’s easiest to stop paying attention.
If AI answers are where buyers’ shortlists start forming, then getting mentioned while a competitor gets recommended means you helped the buyer make sense of the category, and the model pointed them at someone else to solve it.
Your report can prove that you show up, but it can’t prove how often you’re chosen. Without a recommendation rate alongside your share of voice, you won’t know preference is a problem until pipeline slows down, weeks after the shortlist has already closed.
Here’s what that number looks like when you collect it.
The Shortlist Barely Moved. The Recommendation Did.
Gong appeared in 100% of the threads. Clari, Salesloft, and Outreach appeared in nearly all of them. Even as the search paths changed from one person to another, the answers kept returning to the same core group.
The recommendations were spread across that group and beyond it. Gong and Apollo each received three clear recommendations. Two threads leaned toward Salesloft with Clari as a combined approach, and one did the same with HubSpot and Outreach. Two answers didn’t settle on a single vendor.
To separate these outcomes, we counted a vendor as “mentioned” whenever it appeared in an answer, and as “recommended” only when the answer explicitly preferred it or presented it as the best fit. Conditional recommendations and answers with no single preference remained separate.
The image below maps mentions against recommendations across all 11 threads.

This gap isn’t unique to our threads. Niko Alho’s AI Recommendation Index, which analyzed 144 answers across six models, recorded one brand earning 86% share of voice in its category while taking none of the 36 first-pick recommendations.
The heatmap makes the difference hard to miss. Mentions tell you which vendors were repeatedly included in the consideration set, but not why one was preferred. For that, we had to look at how each answer interpreted the buyer’s problem.
The Recommendation Moved With the Job Each Model Appeared To Infer
The models weren’t simply ranking the biggest brand. Their visible reasoning framed the buyer’s problem in different ways, and the recommendation moved with those interpretations.
You can see it happen in one thread. The model’s rationale includes the line, “Gong is great but not ideal for replacing a stack.”
The prompt never said “stack consolidation” should decide the recommendation.
Yet, the answer treated consolidation as the buyer’s main requirement, judged Gong against it, and recommended Salesloft and Clari instead. The rationale suggests the model inferred a requirement the buyer never stated.
Across the 11 threads, four interpretations of the job lined up with four recommendation directions.
When the answer framed the buyer’s job as prospecting, it recommended Apollo. When the answer framed the job as coaching a new group of SDRs, it recommended Gong.
Stack consolidation brought recommendations for Salesloft and Clari, while CRM migration brought HubSpot and Outreach. The prompt never changed, but the job each answer appeared to infer did, and the recommendation moved with it.

Research on underspecified prompts makes that pattern plausible. In a paper published in Findings of ACL 2026, Yang and colleagues found that models inferred requirements users hadn’t specified in 41.1% of cases.
Underspecified prompts were also twice as likely to regress across model or prompt changes. The study wasn’t about software buying, but it helps explain why an open buying question can produce different interpretations of the buyer’s unstated priority.
Rand Fishkin, co-founder of SparkToro, found the same instability at a much larger scale. When he and Patrick O’Donnell ran 2,961 identical prompts across ChatGPT, Claude, and Google’s AI, the odds of the same prompt returning the same list of recommended brands twice were under 1 in 100. As he said:
“Any tool that gives a ‘ranking position in AI’ is full of baloney”
Our 11 threads show a pattern, not a rule. Some answers were conditional, and two produced no single preference. We also couldn’t isolate job interpretation from other variables that may have shaped the outputs.
So, if the models interpreted the buyer’s job differently, why did they keep returning to the same vendors?
The Search Paths Varied. The Same Sources Kept Returning.
The answers rarely relied on a single search. Instead, the prompt opened into several related searches, each exploring a different part of the buyer’s problem. Google calls this pattern “query fan-out” in AI Mode.
Searching widened the question, while interpreting the results shaped what the answer treated as the buyer’s job.
The search paths varied across our threads, but many of the same sources kept returning. Zapier appeared in 70% of the ten participant-level source records, while Oliv, Tellius, and ZoomInfo’s Pipeline each appeared in 60%.
Different routes kept circling back to many of the same roundups, comparison pages, and review sites.

One thread also showed how context outside the prompt can shape the route.
The model brought in a previous conversation about Zoho, read the buyer’s problem as a CRM migration, and shifted its recommendation toward HubSpot and Outreach.
This happened once, so we can’t treat memory as the explanation for the wider pattern. But it shows why the prompt alone may not explain the recommendation.
We also can’t isolate what produced the variation across the other threads. The model, account, memory, settings, or another factor could all have played a role. Our experiment wasn’t designed to separate them, so they remain possible explanations rather than proven causes.
Marketers can’t control the route a model takes through a question, but they can test whether the sources that keep returning associate their brand with a job it credibly performs. How much influence that gives them depends on the question the buyer asks.
What About Branded and Comparison Queries?
When buyers name the vendors in the prompt, the same brand is more likely to remain the leading recommendation across repeated runs.
Conductor found the same pattern at a much larger scale. Across roughly 14,000 API calls, the leading brand held its position in about 91% of comparison-intent prompts, compared with 59% of purchase-intent prompts, the least consistent type it measured.
So a question such as “Gong vs. Apollo” has already narrowed the field, leaving less room for the model to reinterpret what the buyer needs.
Our prompt left that narrowing to the model. It described an active need but named neither the vendors nor the requirement that should decide between them. That places it in the discovery stage, when buyers are still learning the category and forming their initial shortlist.
Discovery Prompts Are Where Recommendations Shift
That boundary is also the opportunity. Before the buyer names any vendors, the recommendation has more room to move with the model’s interpretation of the job.
This is where mention tracking tells you the least, because making the shortlist doesn’t reveal which vendor the answer prefers.
The discovery stage, when a buyer hasn’t named vendors yet, is exactly where mentions are least informative and the model has the most room to steer the recommendation elsewhere. You can’t control which job a model infers from an unbranded prompt, but you can find out whether you’re winning the recommendation or just entering the conversation.
To see whether your brand earns that preference, you need to measure recommendation rate.
How to Measure How Often AI Recommends Your Brand
To measure both outcomes, use the same prompt set for two separate counts.
- Identify a high-value problem customers already choose your product to solve, not a keyword you want to own. Look for that problem in sales calls, win-loss interviews, and customers’ explanations of why they bought your product.
- Collect the unbranded, discovery-stage prompts that express that job: Use problem-shaped questions a buyer would ask before naming vendors, rather than comparison prompts where the field is already narrowed.
- Run each prompt repeatedly across different models, accounts, and weeks. For every answer, record whether your brand was mentioned and whether it was recommended. A single run only shows what the model said once. Repeated runs show whether the same pattern holds.
- Record the recommendation rationale and the sources that recur when you win and lose: Note which job the answer appears to prioritize, which vendor it prefers, why it prefers that vendor, and which sources support the answer. This lets you compare more than the final pick.
- Track your mention rate, recommendation rate, and the percentage-point gap between them over time. These numbers don’t prove ROI, but they give you a stronger leading indicator than share of voice alone. Mention rate shows how often you enter the consideration set. Recommendation rate shows how often the answer prefers you.
Once you can see the gap, you can investigate why the recommendation goes elsewhere.
To Get Recommended, Own One Job in the Sources AI Reads
Our findings point to how clearly recurring sources connect your brand to one job you already win. Identify that problem in sales calls, recent wins, and win-loss interviews, then check whether those sources explain why your product suits it, provide customer proof, and distinguish it from the alternatives.
If the customer evidence isn’t there, the problem sits with the product or positioning, not visibility.
Compare the Sources Behind Your Wins and Losses
Start with the source and recommendation data from your baseline. Separate the answers into two groups: those that recommended your brand and those that preferred another vendor or made no clear recommendation.
Now compare the sources across those groups. Check:
- Which sources recur when your brand wins
- Which sources recur when another vendor wins
- How each source describes your product’s fit for the job
- What evidence it gives for preferring one vendor over another
- Whether your brand appears in the source at all
If your brand rarely appears in the sources that recur, you have a presence gap. If the sources mention you but position another vendor as the better fit, you have a preference gap.
That distinction tells you what to fix. A presence gap calls for earning coverage in the third-party sources that repeatedly surface. A preference gap calls for giving those sources stronger proof that your product fits the job.
Give Third-Party Sources a Stronger Case to Verify
Our analysis of AI answers about B2B software found that nearly 90% of the citations pointed to sources outside the vendor’s own website. So when the same roundups, comparison pages, review sites, and community discussions keep returning, those are the first places to check whether your product is connected to the job you want to own.
Start with the roundups, comparison pages, review sites, and community discussions that appeared most often in your prompt set.
For each source, check whether it contains the evidence a writer would need to describe your product as a strong fit for the job:
- A clear explanation of the use case
- Relevant product capabilities
- Customer proof tied to that use case
- A direct comparison with the alternatives
- Honest limits showing who the product isn’t for
Where the case is missing or outdated, give the publisher evidence tied to the job: current product capabilities, original data, a verified customer result, a direct comparison with alternatives, and the situations where your product isn’t the right fit. The aim isn’t another mention; it’s an accurate account of why a buyer would choose your product for that job.
Compare the Results Before You Scale the Work
After strengthening the job-fit case in the sources, rerun the same prompt set under comparable conditions. Keep the job, models, accounts, and testing period as consistent as possible so you can compare the new results with your baseline.
Look for more than a temporary lift in mentions. The useful pattern is whether your recommendation rate improves for the job you targeted, particularly when the same third-party sources recur behind the answers.
If that pattern holds across repeated runs, expand the test to more prompts or another buyer job. If mentions rise but recommendations don’t, broader visibility hasn’t solved the preference gap. Revisit how those sources explain your fit before investing further.
The comparison won’t prove that one source caused the change. Models, outputs, and source pools continue to vary. It tells you whether the job association and recommendation rate moved together strongly enough to justify a larger test.
Turn AI Mentions Into Recommendations
A healthy share-of-voice score can hide the fact that another vendor is earning the recommendation on the buyer jobs that matter most to your pipeline. The gap between being mentioned and being preferred is where deals move without your team seeing it.
Recommendation rate makes that gap visible. Compare it against share of voice by buyer job, and you can see exactly where visibility becomes preference and where a competitor wins instead.
Foundation’s AI-visibility team can establish that baseline, identify the buyer jobs where your brand loses the recommendation, and map the sources recurring behind those answers.
Book a call to build an AI-visibility strategy focused on earning recommendations, not mentions alone.