web analytics

Three Of Five Prompts Or Stay Home: The New Semrush Study On ChatGPT Category Expansion

SEMrush topic focus study

Every marketing plan eventually reaches the same two questions. Which category do we go after next, and how much do we commit before we know whether it worked?

For organic search those questions have always been answered with judgment and a keyword tool. New research from Semrush, published August 3 and produced with Kevin Indig of Growth Memo, gives them a defensible answer for AI search. Semrush’s research on topical focus and category expansion in ChatGPT is worth reading directly. This article covers what it means for planning, sequencing and budget.

It is the second study in the series. The first, which I analyzed in my write-up of the topic ownership research, established that most categories have no dominant brand and that domain-level SEO metrics predict ownership at roughly the rate of a coin flip. This one takes up the obvious follow-up question: if the usual metrics do not explain who wins a category, what does, and how should a company choose where to compete?

The Research In Brief

The team worked from 1,094 US categories, each represented by five prompt variants, tracked monthly in ChatGPT from January through June 2026. The resulting data set holds 283,215 domain-category citation records, 76,493 observations of a brand being named in an answer, and 45,578 category-expansion appearances mapped across 1,458 brand entities.

Two constructs drive the analysis. A brand’s core expertise was defined as any category where it had already appeared in three or more of the five prompts before the month under measurement. Category closeness was scored by semantic similarity between a target category’s prompts and the prompts of categories the brand already answered well.

The study then asked a direct question: when a brand shows up in a new category, does it get cited as a source, named in the answer, or both, and how does that change as the new category sits further from what the brand already does?

Three Constraints The Data Places On Category Planning

There Is A Minimum Commitment, And It Is Higher Than One Article

Brands present in a single prompt out of the five saw their mention share fall rather than rise. Three prompts was the point at which that penalty disappeared.

That finding removes the most common entry strategy in content marketing. Publishing a single piece into a new category to see whether it gains traction is not a low-cost test in this data. It is an outcome the research associates with going backward on the measure that drives buyer decisions.

The planning consequence is straightforward. Category entry has a minimum viable unit of roughly three of five buyer questions covered with real substance. A budget that funds twelve categories at one asset each and a budget that funds three categories at four assets each cost the same and, on this evidence, produce very different results.

Adjacency Beats Opportunity Size

Proximity changed every outcome measured. Near a brand’s established expertise, 74% of its appearances came with a citation, 44% with a naming in the answer text, and 34% with both. Out at the far end of the similarity range, the same three figures read 50%, 25% and 9%.

The conversion rate sharpens it further. Citations arising in the least related categories carried a named brand 18% of the time. In the most related, that figure reached 46%.

Indig’s summary in the study is worth keeping in front of a planning team:

The categories where a brand consistently earns named mentions are the ones from which it can credibly expand.

For a planning team accustomed to ranking opportunities by search volume, this reorders the list. A large category three steps removed from your established expertise converts citations into mentions at roughly a third the rate of a smaller category next door. The larger prize is discounted by a factor the keyword tool does not display.

Industry Sets A Ceiling You Cannot Publish Past

The pattern was not uniform across sectors. Financial services and real estate converted wider category coverage into citation leverage. The legal and healthcare categories did not: their citation gains from breadth were weaker, and the mention-share trend stayed negative even for brands covering all five prompts.

Companies in regulated, high-stakes fields should read that carefully. Coverage depth appears to be the entry requirement rather than the differentiator. My working hypothesis, offered as hypothesis, is that the lever in those categories is credentialing and third-party validation rather than volume of published material. The researchers did not test author credentials, licensure signals or original research, and said so directly. Anyone in legal or healthcare planning a content investment on this basis should treat the ceiling as real and the explanation as unproven.

A Method For Choosing The Next Category

The study describes its approach clearly enough that an in-house team can reproduce a workable version of it.

Start by identifying anchor categories. These are the categories where your brand already earns both citations and named mentions across most of the buyer question set. Anchors are the only legitimate launch points for expansion, and most companies have fewer of them than they assume.

Score adjacency next. Build the five-question prompt set for each candidate category, embed those prompt sets alongside the prompt sets for your anchor categories, and rank candidates by cosine similarity to the nearest anchor. This is a short piece of Python work using any current embedding model, and it produces a ranked expansion list grounded in the same construct the study used. A team that has never done this before can complete a first version in an afternoon.

Then apply a practical filter to the top of that ranked list. Genuine adjacency usually means the categories share a customer, a problem, a product capability, and a buying process. Where the similarity score is high but those four do not line up, the score is picking up vocabulary rather than market logic.

Sequence the result. Enter one adjacent category at a time, at full three-of-five depth, and let it become an anchor before moving to the next. Expansion in this model compounds outward from established ground. It does not leapfrog.

The Metric I Would Add To Your Reporting

Most AI visibility dashboards report citation count and mention count as separate totals. The study points toward a more useful figure: the rate at which your citations in a category also carry a named brand mention.

Treat that conversion rate as a diagnostic for whether you belong in a category at all. The study’s range gives you the reference points. Around 46% is what the most related categories produced. Around 18% is what the least related produced. A new category where your citations convert at the low end is telling you the audience-facing signal has not followed the content, which usually means the category sits further from your core than the plan assumed.

That single ratio answers a question executives ask constantly and that current reporting handles poorly: is this working, or are we just publishing?

Where The Evidence Runs Out

These relationships are observational. There was no intervention and no control group, and the authors consistently wrote “associated with” rather than “caused by.” That language should survive into your internal decks.

The most substantial open question concerns direction of causation. Thin coverage may suppress mention share. Thin coverage may also be nothing more than what an unfamiliar brand looks like inside a category. Recognition would lag for such a company no matter what it published, which means the single-prompt penalty could be tracking brand size and reporting it as a coverage effect.

Scope is the second limit. Only 1,458 mapped brand entities underpin the expansion analysis, against the 50,000-plus brands behind the first study, and the coverage is limited to ChatGPT in the United States across six months. No data was collected on Gemini, Perplexity, Copilot or Google’s AI Overviews.

The authors also listed what the study does not account for: writing style, content quality, originality, overall brand authority and third-party mentions. That is a substantial list of uncontrolled variables, and it argues for treating the findings as planning guidance rather than as a formula.

What the research does support is a shift in how expansion decisions get made. Category selection driven by proximity to demonstrated expertise, entry funded at a minimum depth rather than as a single test asset, sector-specific expectations for what content alone can achieve, and reporting that separates the sources a system links from the brands a system names. Those four changes are defensible on this evidence, and they cost nothing beyond the discipline to sequence work that most organizations would rather run in parallel.

Scroll to Top