Home/Blog/How many prompts should I track?
PromptDNA

How many prompts should I track? The honest math

Every LLM visibility vendor prices by prompt volume — and every sales conversation nudges you toward the bigger package. But most buyers are paying to generate data that will never be read by anyone. The right prompt count isn't a package tier; it's a function of which of two very different jobs you're hiring the data to do.

How many prompts do you actually need to track?

Fewer than the pricing page wants you to buy. For content-gap insights — the use case where LLM response data actually changes what you publish — 100 to 250 intent-mapped prompts on one or two engines is enough for most businesses. For pure market-share monitoring across many engines, 300 to 1,000 prompts make sense, but on a monthly cadence, not daily. The bill is driven by a multiplier — prompts × engines × iterations × frequency — and cutting the factors nobody looks at (engines, frequency) saves far more than cutting prompts. If your monthly report consumes one percentage number, you don't need a firehose to produce it.

Why prompt count — not features — drives your bill

LLM visibility tools price almost universally by prompt volume, because every tracked prompt costs the vendor real inference money. A sample of published pricing in 2026:

ToolEntry tierWhat you getEffective price per prompt/mo
Semrush AI Toolkit$99/mo per project25 prompts~$4.00
Otterly Lite$29/mo15 prompts~$1.93
Otterly Premium$489/mo400 prompts~$1.22
Profound Starter$99/mo50 prompts, 1 engine~$2.00
Profound Growth$399/mo100 prompts, 3 engines~$4.00
Adobe LLM Optimizerprompt packages1,000-prompt minimum

Notice two things. First, the per-prompt price drops with volume — the classic upsell architecture: the bigger package always looks like the better deal per unit. Second, minimums are creeping up; Adobe won't sell you fewer than a thousand prompts at all. The vendor's incentive is to size your package by what generates their revenue, not by what your reporting consumes. Nobody on a sales call has ever been told they need fewer prompts.

The uncomfortable question to ask before any purchase: which report, read by which person, consumes this data? Every prompt that doesn't feed an answer to that question is budget converted into database rows.

The multiplier nobody puts on the pricing page

The prompt count is only one of four factors. Your actual data volume — and with usage-priced tools, your actual cost — is:

prompts × engines × iterations × frequency

Iterations matter because a single LLM response is an anecdote: the same prompt asked five times returns different answers, so serious tracking runs each prompt several times and averages. Now watch the multiplier work:

  • The "big package" setup: 500 prompts × 5 engines × 3 iterations × daily = 225,000 responses a month.
  • A focused setup: 150 prompts × 1 engine × 3 iterations × weekly = ~1,800 responses a month.

That's a 125× difference in generated data — and in most organizations, both setups feed the same monthly slide with one visibility percentage on it. LLM answers also don't change meaningfully day to day for most brands; models update in cycles, and content changes take weeks to propagate. Daily tracking of a stable brand mostly measures noise, expensively.

Job 1: LLM market-share monitoring

The first legitimate use of visibility data is the panorama: what share of AI answers mentions us, across ChatGPT, Gemini, Claude, Perplexity, DeepSeek, Kimi — and how does that compare to competitors? A defined prompt set, say 500, runs either once ("where are we today?") or on a recurring cadence, and the output is a share-of-voice KPI, ideally benchmarked against the competition.

This is a real product with a real audience — brand teams and executives want that number, and agencies need it for reporting. But be honest about its data economics: the thousands of stored responses are consumed as one aggregated percentage. Almost nobody reads the responses themselves. That has two consequences for sizing:

  • Sample, don't census. Share of voice is a statistical estimate; 300 well-distributed prompts estimate it nearly as well as 1,500. Precision beyond the decision it informs is waste.
  • Cadence beats volume. A monthly run tracks every trend an executive dashboard can act on. Reserve weekly runs for periods where something is actively changing — a campaign, a migration, a competitor launch.

Job 2: response mining to close content gaps

The second job is smaller in volume and much bigger in value: collect response data from one or two major LLMs and actually read it. Visibility is still measured — but the responses themselves become the working material.

The pattern: your data shows that for a business-relevant prompt — say, "best warehouse management software for mid-size logistics companies" — your client appears in no LLM answer, while two competitors are recommended every time. That's not just a KPI gap; it's a diagnosis waiting to be read. Open the actual responses and you see what the engines say instead, which sources they cite, and what shape that content has: comparison tables, pricing clarity, integration lists, review citations. Now the client knows exactly what to build, and why their current page doesn't match the question. Publish, wait a cycle, re-run the same prompts, and watch whether the answer changes — a closed feedback loop between LLM data and the content roadmap.

Two engines are enough for this job because content strategy is not written per-LLM: the structural work that wins a citation in ChatGPT — clear entities, citable passages, answer-shaped pages — carries over to Gemini and Perplexity, since every engine solves the same retrieval problem. Collecting the same insight five times doesn't make it five times more actionable; it makes it five times more expensive.

Right-size with funnel stages: the geobubbles method

The fastest way to shrink a prompt set without losing signal is to stop treating all prompts as equal. In geobubbles, every prompt is categorized by funnel stage — TOFU, MOFU, BOFU — and the budget conversation changes immediately:

  • BOFU (bottom of funnel) — "is [brand] worth it?", "[brand] pricing", "alternatives to [competitor]". Purchase-ready questions, closest to revenue, and the smallest set: most businesses have 30–60 that matter. Track these first, always, on every plan.
  • MOFU (middle) — comparisons and evaluations: "[brand] vs [competitor]", "best [category] for [audience]". Where recommendations are won or lost; typically 50–100 prompts.
  • TOFU (top) — broad category and how-to questions. Highest volume, weakest link to revenue, and the first place to cut. A thin slice of 20–40 representative prompts tracks the trend; two hundred TOFU prompts track the same trend at ten times the cost.

A 60/30/10 budget weighted toward BOFU+MOFU beats a 10/30/60 set of the same total size on every metric a client pays for — because when a BOFU answer changes, revenue is on the line, and when a TOFU answer changes, a chart moves. Intent mapping is also what kills the classic waste pattern: the generic 500-prompt list that feels thorough and measures nothing. (Why those lists happen and how they burn budgets deserves its own article — it's coming.)

Three setups that work

Small business / single brand: 50–100 intent-mapped prompts, ChatGPT only, weekly or bi-weekly, 3 iterations. Enough to know where you stand on everything that touches revenue, and to mine responses for content gaps. This fits entry-level pricing almost everywhere — you don't need the mid tier.

Mid-market in-house team: 150–250 prompts on ChatGPT + one secondary engine for the insight work, plus a broader 300–500-prompt market-share snapshot run monthly across the engine panorama for the executive dashboard. Two jobs, two configurations, one tool.

Agency, per client: start every client at the small-business setup and let the data argue for expansion. The monitoring line item stays affordable, margins stay healthy, and when a client asks "shouldn't we track more?", you have response data — not a vendor's pricing page — to answer with. In geobubbles, both modes are configured per client and per project, so right-sizing is a settings decision, not a re-platforming.

When a big package IS the right call

Honesty cuts both ways — sometimes the large tier is justified: multi-market brands tracking prompts per language and country (the multiplier works against you legitimately there); enterprises in fast-moving categories where answers genuinely shift week to week; portfolio companies monitoring many brands; and agencies aggregating dozens of clients under one contract. The test is always the same: every prompt maps to a decision someone will actually make. If you can't name the decision, don't buy the row in the database.

FAQ

Q: How many prompts should a small business track?
A: 50–100 intent-mapped prompts on one engine (usually ChatGPT), run weekly or bi-weekly with a few iterations, covers everything revenue-relevant for most small businesses — and fits entry-tier pricing. Grow the set only when the data shows a reason.

Q: Do I need to track every LLM?
A: No. For content-strategy work, one or two engines are enough — the fixes that improve visibility in ChatGPT carry over to Gemini, Claude and Perplexity. Broad multi-engine tracking is a reporting product for stakeholders who want the market panorama; run it monthly, not daily.

Q: Is daily tracking worth it?
A: Rarely. LLM answers for stable brands change in cycles, not days — daily runs mostly store noise. Monthly tracks every trend an executive dashboard can act on; switch to weekly temporarily when something is actively changing.

Q: Why do vendors push high prompt packages?
A: Every tracked prompt costs the vendor inference money and earns them margin — pricing pages are built so the bigger tier always looks cheaper per prompt, and minimums keep rising. The fix is to size from your reporting needs backward: which report, read by whom, consumes this data?

PromptDNAPrompt trackingLLM visibilityCost controlTOFU MOFU BOFUMonitoring
JW

Jonas Weber

Agency Success Lead. Spent 15 years selling and delivering technical SEO consulting before joining Bubbles1. Helps agencies package audits, workshops and retainers that clients actually renew.

Track what pays, not what's pitched

PromptDNA maps every prompt to funnel intent, so your budget goes into answers someone acts on. Free for 14 days.

14-day trial · No credit card · Cancel anytime