How much a different 25 changes the leader
Guides on prompt tracking recommend anywhere from 25 to 50 prompts (Averi) to between 50 and a few hundred (cloro). To see what the choice of questions does at those sizes, we treated one week of answers as a fixed pool and drew 10 or 25 questions from it, 10,000 times each. Each draw was ranked by points for each brand’s place in each answer, an approximation of how the category ranking is scored, and compared with the ranking of the full question set.
DecaGEO, ChatGPT (GPT-5.6 Terra), US, English, answers from the week of October 4, 2026. Share of 10,000 random draws without repeats. Definitions in How we measured.
Going from 10 to 25 questions raised the leader’s hold from 34.0% to 46.6% in CRM and from 39.6% to 52.6% in Marketing Automation. At 25 questions the full-set leader still lost first place in about half the draws in both. In CRM that week, another brand was recommended in more answers than the leader. The leader still ranked first because the ranking also counts how near the top of each answer a brand appears (see How we measured). In Marketing Automation the second most recommended brand trailed the leader by a few answers. In AI Image Generators the leader held first place in 92.0% of 10-question draws. The brands below it moved more. Its second and third brands were nearly tied on that week’s ranking.
The same draws moved the leader’s share of answers, the share of answers that recommended it. In CRM, 49.4% of the full week’s answers recommended the leader, and the middle 90% of 25-question draws put that share anywhere from 36% to 64%. In Marketing Automation the full-set figure was 47.5% and the middle 90% of draws ran from 36% to 60%. In AI Image Generators it was 79.7%, and 25-question draws ran from 68% to 88%.
Share of answers recommending the category leader. Bars show the middle 90% of 10,000 random draws of 10 or 25 questions, and the tick marks the full question set. DecaGEO, ChatGPT (GPT-5.6 Terra), US, English, week of October 4, 2026.
Asking the same question again changes less
Over the four weeks from September 13 to October 4, 2026, every question in all ten categories was asked once a week under the same model and the same question set. For each category’s leader, we recorded whether each answer recommended it. Two numbers come out of that. One is how far a question’s four-week recommendation rate sits from the other questions’ rates (variance between questions, 0.186). The other is how much a single question’s answer flipped from week to week (variance within a question, 0.047). The first was about four times the second. In 77.5% of questions, the leader was either recommended in all four weeks or in none of them. The gap held in AI Image Generators too, where the leader kept first place in almost every draw. Its variance between questions was 0.125 against 0.024 within a question. cloro’s sample-size guide reaches the same advice from statistics. Its reasoning is that “variance between prompts is larger than variance within one, so breadth buys more statistical information than depth.” In these ChatGPT answers the difference was about fourfold. The draws came from one question set, written from buyer personas for each category. Questions a team writes by hand can differ from each other more than questions from one generated set, so the spread for a hand-picked 25 could be wider. That has not been measured.DecaGEO’s rankings are a value of their question set too
The CRM leader in the week of October 4 is first across DecaGEO’s full CRM question set, and a 25-question slice of that same set named a different leader in 53.4% of draws. Each category’s set is built from buyer personas that combine the decision conditions buyers state, such as business size, primary use case or budget priority, and no question names a brand. The set stays the same from week to week, so a weekly move compares the same questions. Once you sign in, the Prompt & AI Response tab shows how many questions the category’s set had that week.What to do with your own prompt set
- Decide which buyer situations to cover before deciding how many prompts. List the conditions buyers in your category state when they ask for a recommendation, such as team size, main use case, required integrations and budget, and give each condition questions of its own. cloro’s guide puts it as “every distinct buyer intent you care about.” The count follows from that list.
- When budget forces a choice, add questions before adding reruns. In DecaGEO’s answers, a different question changed whether the leader was recommended about four times as much as the same question a week later.
- Mark every week your prompt set changed. A number from a new set is a new baseline. In a report, start a new line at that week instead of connecting it to the old one.
- Check how much your own leader depends on your set. If your tool exports per-prompt results, split your prompts at random into two halves in a spreadsheet. In each half, rank brands by how many prompts recommended them and see whether both halves put the same brand first. Repeat with a few different splits. If the halves name different first brands, report the top brands as a group, not a single winner.
- Use a fixed set to find buyer conditions to cover. DecaGEO’s public category rankings show each week’s leader from the full persona-based question set without signup. Once you sign in, including on the free plan, the Segment Position tab shows your rank among the answers to each buyer condition’s questions, with “Not Mentioned” where ChatGPT did not recommend you, and the week selector covers the latest four weeks. A condition’s result comes from that condition’s questions only, so use the conditions to choose the buyer situations to write questions for in your own set, not as a ranking of their own. Look across all four weeks so that a single week’s flip is not read as a gap. Prompt & AI Response lists the questions whose answers mentioned your product, with their persona tags, which shows how buyers under each condition phrase the question. Pro adds full answer text and CSV export.
FAQ
How many prompts should I track for AI visibility?
Enough to give every buyer situation you care about its own questions. In DecaGEO’s CRM and Marketing Automation answers for the week of October 4, 2026, 25 random questions kept the full set’s leader first in 46.6% and 52.6% of draws. In AI Image Generators, where the leader was far ahead, the figure was 99.98%. See How much a different 25 changes the leader.Is it better to run more prompts or rerun the same prompts more often?
More prompts. Across ten categories from September 13 to October 4, 2026, ChatGPT’s recommendation of each category’s leader varied about four times as much between questions as within one question from week to week. See Asking the same question again changes less.My share of voice jumped the week we changed our prompts. Is that real?
Not as a trend. A number from a new prompt set starts a new baseline. In the CRM answers, 25-question samples from a single week put the leader’s share anywhere from 36% to 64%. See What to do with your own prompt set.Does this apply to DecaGEO’s own rankings?
Yes. Each ranking is a value of that category’s full question set. The set is built from buyer personas and stays the same every week, so weekly moves compare the same questions. See DecaGEO’s rankings are a value of their question set too.How we measured
The panel’s questions, schedule and metric definitions are on the methodology page. What is specific to these figures:- Window. Weeks of September 13 to October 4, 2026, all measured on GPT-5.6 Terra with the same question set in each category. The draws use the week of October 4 only.
- Leader. The brand ranked first in the category for the week of October 4. For the draws, the full set was re-scored the same way as each draw and gave the same first-place brand in all three categories.
- Scoring a draw. Only brands on that week’s category ranking count. In each answer, a brand earns more points the nearer the top of the answer’s list it appears, and a draw ranks brands by their total points. This approximates how the DECA Score is built. “Same top 5” compares the five highest totals with the full set’s five.
- Draws. 10,000 random sets of 10 or 25 questions per category, each question at most once per set. The range is the 5th to 95th percentile of the leader’s share across draws.
- Share of answers. The share of the week’s answers in which the leader was among the recommended brands.
- Variance. For each category’s leader and each question, the share of the four weeks in which the answer recommended it. Variance within a question is the average of p(1−p) over questions. Variance between questions is the variance of p across questions in the category. The ten-category figures average the categories, weighting each by its number of questions. Each question is asked once a week, so “again” means a week later.

