How to Measure Brand Sentiment in AI-Generated Content

How to Measure Brand Sentiment in AI-Generated Content
Most AI visibility programs stop at counting. They track how often a brand gets named in ChatGPT, Perplexity, or Google's AI Overviews, chart the mention rate over time, and treat a rising line as progress. That count is the easier half of the problem, and by now it is a reasonably well-served one. Plenty of tools will tell you whether your name showed up.
What almost nobody measures well is how the model described you once it did. A brand can be mentioned in nine of ten answers and framed in every one of them as the expensive option, the legacy choice, or the tool people outgrow. The count looks healthy. The narrative is doing damage.
At Geostar we run GEO programs where that gap between mention volume and mention quality shows up constantly, and the measurement side of it is genuinely unsettled. Below we cover why sentiment behaves differently from visibility, why a single query tells you almost nothing, the sycophancy problem that breaks conventional sentiment scoring on AI text, a practical measurement framework, the mistakes that produce misleading data, and what to do once the numbers come back.
Why Sentiment Is a Different Problem Than Visibility
Visibility tracking asks a binary question: did the brand appear in the response. Sentiment tracking asks a qualitative one: when it appeared, what was said about it. The first is countable and reproducible enough to trend. The second requires interpreting language that changes wording every time you ask, which is why most measurement stacks quietly skip it.
In AI answers the distinction carries far more weight than it did in traditional search, because a synthesized response characterizes you rather than merely listing you. A blue link on a results page carried no editorial opinion. A synthesized paragraph naming three vendors and calling one of them "better suited to enterprise teams with dedicated support budgets" has delivered a verdict, and the buyer reads that verdict rather than clicking through to check it.
Even the countable half is noisier than most teams assume. SparkToro and Gumshoe ran the largest study on this to date in January 2026, with 600 volunteers running 12 prompts a combined 2,961 times across ChatGPT, Claude and Google's AI. They found there is less than a 1-in-100 chance that ChatGPT or Google's AI will return the identical list of brands twice for the same prompt, and roughly 1-in-1,000 odds of returning the same order twice [1]. Their City of Hope example makes the practical implication concrete: the hospital appeared in 69 of 71 ChatGPT responses for a category query, a 97% visibility rate, yet was the top-ranked mention in only 25 of them.
The researchers who ran that study did not attempt to quantify sentiment at all. They flagged it as an open problem, harder than list consistency, and left it there. That is the honest state of this field right now. If brand-list reproducibility is the easy version of the question and it still required 2,961 runs to characterize, the qualitative layer sitting on top of it deserves more methodological care than a weekly spot-check. Our framework for analyzing AI-driven brand mentions covers the mention-tracking foundation this sits on.
Why a Single AI Query Never Tells You the Real Sentiment
The instinct when someone asks how AI describes your brand is to open ChatGPT, type the question, and read the answer. That read is not a measurement. It is one draw from a distribution.
Harvard Business School researchers studying LLMs for market research treat this as the defining property of the output. Their working paper on the subject queries models dozens of times per survey question specifically because responses form a distribution rather than a single deterministic answer, and they found that willingness-to-pay estimates derived from those distributions are sometimes comparable to human study results but often inaccurate and in some cases wrong-signed [2]. If a carefully constructed academic sampling procedure still lands on the wrong sign occasionally, a one-off query read at a desk carries no evidentiary weight at all.
The problem compounds because the instrument has the same defect as the object. Teams that automate sentiment classification usually hand the AI response to another LLM and ask it to label the tone. A peer-reviewed review of what the literature calls the Model Variability Problem documents exactly what goes wrong there: LLMs used as sentiment classifiers produce inconsistent classifications of the same text, both across models and across repeated runs of the same model [3]. So the thing being measured varies run to run, and the tool doing the measuring varies run to run on top of it.
Practically, that means two different measurement postures:
Neither the object nor the instrument is stable enough to trust once. Sampling is not a refinement here. It is the minimum condition for the number to mean anything.
The Sycophancy Problem: Why AI Text Can Look More Positive Than It Is
There is a bias in AI-generated text that has no equivalent in the human-written social posts and reviews traditional sentiment tools were built for, and it systematically pushes readings toward the positive.
A Stanford-reported study published in Science in March 2026 tested 11 models and found they endorsed a user's stated position 49% more often than human respondents did, and endorsed described harmful or illegal behavior 47% of the time [4]. Models are agreeable by default. That tendency does not switch off when the subject shifts from a personal dilemma to a brand.
The mechanism is what turns this into a measurement problem instead of a curiosity. The excess agreement is not delivered as obvious flattery. The Stanford reporting describes it arriving in seemingly neutral, academic-sounding language, with one example framing clearly wrong behavior as stemming "from a genuine desire to understand the true dynamics." There are no exclamation points and no superlatives. The text reads measured while the substance tilts.
Traditional sentiment scoring cannot see this. The classic method used across the social-listening tools that dominate this topic assigns each mention a positive, neutral, or negative label, often via keyword or lexicon matching, then sums the labels into a score. Feed that method a sycophantic AI paragraph and it returns "neutral," because the vocabulary is neutral. The framing bias passes through the scorer untouched, and the resulting dashboard understates how much of your apparent neutrality is really model-default agreeableness rather than earned positioning.
Two corrections follow. First, sentiment classification on AI text has to read framing and recommendation strength, not vocabulary: whether the model recommends you, hedges on you, or positions you as the fallback. Second, absolute sentiment scores are close to meaningless in isolation. What carries signal is your framing relative to named competitors on the identical prompt set, since the sycophancy bias applies to all of you at once and cancels out of the comparison.
A Practical Framework for Measuring AI Sentiment
Practitioner methodology has converged on a repeatable procedure. Six steps, each addressing one of the failure modes above.
- Build a standardized prompt set. Oltre.ai's methodology recommends 30 to 50 standardized prompts per product line or topic as a reliable baseline [5]. Span informational queries ("what is X"), comparison queries ("X vs Y"), and decision-stage queries ("best X for enterprise teams"), because sentiment shows up most sharply in the decision-stage answers where the model actually picks.
- Run every prompt across multiple platforms. The major assistants and AI Overviews behave differently enough that a single-platform reading is a platform artifact, not a market signal. Oltre.ai notes that ChatGPT and Claude tend to synthesize opinionated recommendations while Gemini and Perplexity anchor more tightly to cited web entities, which changes both what gets said about you and where the framing originates.
- Store the full response and an evidence snippet for every mention. Oltre.ai flags skipping this as the biggest mistake teams make [5], and we agree. A sentiment score with no retrievable text behind it cannot be audited, disputed, or diagnosed. When a score moves, the snippet is the only thing that tells you why.
- Classify at the aspect level, not the response level. A single answer routinely praises your onboarding and criticizes your pricing. Collapsing that into one blanket label discards the actionable half. Tag the theme alongside the polarity so the output tells you which specific narrative needs work.
- Re-run the same prompt set two to three times in a short window. Siftly recommends this cadence to establish stability, and their worked example is a useful calibration: a brand appearing in 7 of 10 runs represents roughly a 70% mention rate with approximately plus or minus 15% confidence at a 95% confidence interval [6]. That error bar is wide, and knowing it is wide stops teams from reacting to noise.
- Benchmark against three to five named competitors and trend it. Run competitors through the identical prompt set in the same window. The comparison neutralizes the sycophancy floor and turns an uninterpretable absolute score into a position.
What to actually track, once the pipeline is running:
Most teams find the pipeline itself is not the hard part. Interpreting the output against a strategy is, and our guide to the impact of AI on brand visibility covers how these readings fit the wider picture. If you would rather see your own numbers before building any of this, book a free audit and we will run the prompt set for you.
Common Mistakes That Produce Misleading Sentiment Data
The failure modes are consistent across the teams we audit:
- Trusting a single query as a valid read. The HBS, Model Variability, and SparkToro findings all converge here from different directions [1][2][3]. One query is an anecdote with a number attached to it.
- Treating raw mention volume as the headline. High-frequency mentions framed negatively are worse than low-frequency mentions framed as the recommendation. Share of voice without sentiment weighting can point a team confidently in the wrong direction.
- Running keyword or lexicon scoring on AI text. As the Stanford work shows, sycophantic positivity arrives in neutral vocabulary, so lexicon methods systematically misread it.
- Reading order as meaning. SparkToro's data puts the ordering of brands at roughly 1,000-to-1 odds against reproducing, against under 100-to-1 for the list itself. Position within an answer is the least stable thing on the page.
- Skipping evidence snippets. Without stored text, every label is a black box, and nobody can tell a genuine narrative shift from a classifier having an off day.
One further note on sourcing. Oltre.ai reports a meta-analysis finding that users click citations inside AI summaries at roughly 15 times lower rates than traditional search links. We have not been able to trace that figure to its primary study, so treat it as reported rather than established. The directional point behind it, that in-answer framing matters more when fewer people click through to verify it, holds regardless of the exact multiple.
What to Do With the Data
Sentiment findings are only worth collecting if they change what gets published. Three moves cover most of what the data will surface.
Correct inaccurate third-party content first. Source attribution usually reveals that a specific outdated review, comparison page, or forum thread is supplying the framing the model repeats. Fixing the source is more durable than trying to out-publish it, and it addresses the problem where it lives.
Refresh owned content that is feeding stale framing next. Models pull heavily from what you have already published, so an old pricing page, a deprecated positioning statement, or a feature list you have since outgrown will keep resurfacing in answers long after the company moved on.
Then produce content that directly addresses the recurring negative themes. If aspect-level data shows the framing consistently turns negative on implementation complexity, the answer is a substantive, well-sourced piece on implementation, not a positioning tweak.
The commercial stakes justify the work. Semrush's study of AI search traffic found that the average LLM-referred visitor is worth 4.4 times the average traditional organic search visitor [7], and their analysis explicitly recommends managing negative online sentiment as a core AI-visibility lever. Traffic arriving from AI answers converts because the model has already done the recommending. That advantage only holds while the recommendation is favorable. Our complete guide to generative engine optimization covers the broader program these fixes belong to.
Frequently Asked Questions
How is measuring sentiment in AI-generated content different from traditional social sentiment analysis?
Traditional sentiment analysis reads human-written text that stays fixed once published, so one pass over a corpus gives you a stable answer. AI-generated text is regenerated on every query and varies substantially between runs, which means you are measuring a distribution, not a document. AI text also carries a systematic agreeableness bias that human text does not.
How many times do I need to query an AI platform before I can trust the sentiment reading?
Practitioner guidance lands around 30 to 50 prompts per topic, re-run two to three times in a short window [6]. Siftly's calibration example is worth internalizing: a 7-of-10 result gives roughly a 70% rate at plus or minus 15% confidence, so anything under about 10 runs per prompt produces an error bar too wide to act on.
Can keyword-based sentiment tools accurately score AI-generated text?
Not reliably. Lexicon scoring keys on vocabulary, and the Stanford research shows that AI's excess positivity arrives in neutral, academic-sounding language rather than obvious praise [4]. Those tools will return "neutral" for text that is meaningfully tilted, so you need classification that reads framing and recommendation strength instead.
How often should I re-measure AI sentiment for my brand?
Monthly is the common practitioner cadence, and it fits the pace at which models update and narratives drift. Move to weekly around a launch, a pricing change, a competitor's funding announcement, or any event likely to generate a wave of new third-party content the models will ingest.
What do I do if AI platforms are describing my brand negatively?
Start with source attribution to find where the framing originates, since it is usually traceable to a handful of specific pages rather than diffuse model opinion. Correct inaccurate third-party content, refresh outdated owned pages, and publish substantive material addressing the recurring theme. If you want a read on your current framing and where it is coming from, book a free audit and we will map it against your competitors.
References
[1] Rand Fishkin, with Patrick O'Donnell (Gumshoe.ai). "NEW Research: AIs are highly inconsistent when recommending brands or products." SparkToro, 2026. https://sparktoro.com/blog/new-research-ais-are-highly-inconsistent-when-recommending-brands-or-products-marketers-should-take-care-when-tracking-ai-visibility/
[2] James Brand, Ayelet Israeli, and Donald Ngwe. "Using LLMs for Market Research (Working Paper 23-062)." Harvard Business School, 2026. https://www.hbs.edu/ris/Publication%20Files/23-062_1f58623a-ee21-44b9-a262-276047bc5543.pdf
[3] D. Herrera-Poyatos et al. "An overview of model uncertainty and variability in LLM-based sentiment analysis." PMC / NCBI, 2025. https://pmc.ncbi.nlm.nih.gov/articles/PMC12375657/
[4] Ula Chrobak. "AI overly affirms users asking for personal advice." Stanford News, 2026. https://news.stanford.edu/stories/2026/03/ai-advice-sycophantic-models-research
[5] Oltre AI. "How to Track Brand Sentiment in LLMs: A Complete Analysis of AI Citation Quality." Oltre.ai, 2026. https://www.oltre.ai/blog/tracking-brand-sentiment-in-llms-ai-citation-quality/
[6] Uma Maheswari. "How to Track Brand Mentions Across AI Platforms." Siftly.ai, 2026. https://siftly.ai/blog/track-brand-mentions-ai-platforms-chatgpt-perplexity
[7] Rachel Handley. "We Studied the Impact of AI Search on SEO Traffic. Here's What We Learned." Semrush, 2025. https://www.semrush.com/blog/ai-search-seo-traffic-study/
Your Brand is Invisible to AI
While you're optimizing for Google, your competitors are winning in ChatGPT, Claude, and Perplexity. Get the AI visibility analytics you need to drive real inbound growth.
Join marketing teams already winning in AI search.