Citation measurement protocol
How we measure whether an answer engine cites a source
Cited·Index uses on-demand API sampling to record answers and source annotations for questions about the tools in this index and their category. Recurring collection is not active. This page describes the collector as built and the limits of interpreting its historical observations; it is not a measurement of current consumer-app behaviour.
On this page
What counts as a citation
A source annotation is a URL or title returned with the API answer. We keep those annotations separately from the answer prose. A directory-source hit means the collector matched this directory's domain in an annotation URL or title; it does not mean that the tool named in the question was mentioned, recommended or cited on its own website.
A product mention in answer text, a directory page in the source list, and evidence supporting a particular claim are different observations. The collector does not score product mentions or check whether a cited page supports an answer's claims. An annotation alone proves neither that the page was retrieved during this request nor that it earned traffic.
A redirector is not an identified publisher. For known redirector URLs, host attribution uses the annotation title when it looks like a hostname. If the host cannot be established, it remains unknown; the URL and title are retained, not discarded. Directory matching is string-based and host attribution is best-effort, so both can need review.
API configurations and requested models
The collector supports four provider routes through DataForSEO's metered API. These labels name API configurations, not browser sessions in ChatGPT, Claude, Perplexity or Gemini. An on-demand run can select only some of them. The currently requested model aliases are:
- ChatGPT API configuration:
gpt-4o-mini - Claude API configuration:
claude-haiku-4-5 - Perplexity API configuration:
sonar - Gemini API configuration:
gemini-2.5-flash
The request sends the question as user_prompt, the model alias as model_name, and web_search: true. That flag requests web search; the saved response does not establish whether search ran, which retrieval steps ran, or which tool configuration the provider ultimately used.
The collector does not set or verify a controlled country, locale or language for these answers. It does not reproduce a consumer account's location, subscription, personalisation, memory or interface. The aliases above are not immutable model versions: a provider can change the model or retrieval behaviour behind an alias without the saved label changing.
How the question set is selected
The collector reads a tracked-question file; it does not regenerate that file itself. A separate generator combines published-listing warehouse snapshots, live collection sitemaps, retrieval logs, editorial head terms and a pinned category panel chosen from external search-demand data. A stale file can therefore describe an earlier corpus. Publishing or retiring a page does not by itself change a question already in that file.
The generator labels questions with five tiers:
- Vendor. An alternatives question for each tool marked published in the warehouse snapshot.
- Vendor (hot). The vendor questions restricted to listings selected from recent engine retrieval logs. This is a subset of our own corpus, not general buyer demand.
- Filtered. A question derived from each collection-page slug in the sitemaps.
- Head. A small, editorially chosen list of broad category questions.
- Category. A frozen panel selected from external keyword-demand data rather than our page slugs. The generator reuses the pin until a deliberate re-pin changes it; ordinary question-file generation does not select a fresh panel.
Category selection applies an in-category vocabulary and an excluded-industry vocabulary, then collapses some near-duplicates before taking the highest-volume eligible terms. These are editorial sampling rules, not a representative sample of all questions people ask. The chosen questions, tier, model and observation date bound any finding; a category label alone does not support a category-wide conclusion.
On-demand sampling and observation dates
No recurring citation collection is active. All configured tier sets are unarmed. Their retained weekly and monthly calendar slots do not mean those passes run. Measurements require an explicit on-demand invocation; there is no promised next pass, publication cadence or freshness SLA.
The collector attempts one request per selected query entry and API configuration in a pass, not a controlled series of repeated trials. Vendor and hot tiers can contain the same question, so selecting both can repeat its text. A single returned answer is one observation, not a stable citation probability or proof of a trend.
Current rows carry the UTC date assigned once for the run, the question, tier, API route label and requested model alias. This is not an exact per-request timestamp or a pinned model version. Historical observations describe their recorded date and configuration, not what a model or consumer app says today.
Saved answers, missing attempts and partial runs
Successfully parsed responses are kept even when no source annotations are returned. An empty returned source list is not the same as a failed or unasked query, and it does not prove that the answer had no underlying sources. A directory-source rate can use only the observed answers in its declared sample; missing attempts must not be turned into zeros.
A cap abort preserves the successes already collected. A per-pass call cap or provider spend/balance refusal stops the remaining questions and any later sets. If results exist, they are saved and the invocation exits non-zero. If no results exist, no result file is written. Partial observations are not discarded and must not be presented as a complete pass.
Other per-query exceptions are logged and collection continues without a result row for that attempt. The process can therefore finish successfully with questions missing. Cap status is reported in logs and the exit code, not as a structured completion record in the result file; there is no saved manifest of every attempted, failed or unasked query. A file's existence or a successful exit is not proof of a complete edition.
Result files are keyed by date. A later result for the same tenant, question and API route on that date replaces the earlier one, even if the tier or requested model changed. The files are not an immutable archive of every call. The current collector retains answer text and source annotations; older observations may lack answer text or metadata added later.
How this can be wrong
These limits apply even when a request succeeds:
- Single-shot sampling. Answers and source lists vary. Differences between observations can reflect stochastic variation, changed questions, models or collection coverage; the current sampling does not isolate those causes.
- Corpus and query selection bias. Vendor, hot and filtered questions follow our corpus. Head terms are editorial choices, and category questions follow a bounded keyword-demand panel. None is a probability sample of all buyers, prompts or markets.
- A dated category boundary. Vocabulary filters can omit relevant questions or admit ambiguous ones. Demand recorded when a panel was pinned is not current demand, and keyword popularity does not establish how often a question is asked in an AI product.
- API-model scope. Results belong to the requested configuration, not every model sold under the same brand and not the current consumer interface.
- Uncontrolled context. Locale, native tool settings, retrieval execution and consumer account state are not established by these records. An English question does not establish an English-language or US-controlled measurement.
- No exact replay guarantee. The saved question, date and model alias are a locator, not a full native reproducibility context. We do not retain a complete native request/response trace, exact model revision, seed, sampling settings or retrieval trace. Model-version changes may be invisible behind an unchanged alias. Repeating the question cannot recover the earlier answer or prove that conditions were identical.
- Annotations are not verified claims. Host attribution can be unknown or wrong. The collector does not follow each citation, assess its support for a sentence, or measure product mentions. Missing annotations and missing requests have different meanings.
Interpreting and correcting a published reading
A dated edition must be read with its own question set, observed sample and known omissions. If its conclusions assume a complete pass, comparable model versions or verified source hosts that the records do not establish, those conclusions need qualification or correction. A new run would be a new dated observation, not a replacement for the historical answer. This protocol does not promise automatic re-runs, a correction deadline or a complete edition from every invocation.
The data
No edition or downloadable dataset is linked from this page. That does not mean no historical observations exist; it means this page is a method description, not a complete public archive.
Missing dates are unobserved, not zero-citation days. Historical answers cannot be back-filled by asking today's model, and changing this method description does not refresh their dates.
Independence
Cited·Index sells placement and never a verdict, and this measurement is not for sale in either direction: no listed party can pay to be measured, to be excluded, or to have a reading changed. Where a tool we measure is also a commercial relationship, that is disclosed on the page where it appears.
This is a different document from how we rank, which covers how a tool earns a page here. This one covers only how the citation measurement is taken. Found an error in the method? Tell us — a protocol that cannot be corrected in public is not a protocol.