Research

Research reportgeo

Which AI-Visibility Tools Show Their Work? A Methodology Census

We checked 74 AI-visibility tools for a public methodology behind their accuracy claims. 36 claim precision with zero evidence. Only 6 show their work.

Every AI-visibility tool sells the same underlying promise: trust our number for how often your brand shows up in ChatGPT, Perplexity or Google’s AI Overviews. That number is usually the entire product. So a fair question to ask before buying one is the question these tools ask of AI engines all day: show your work. We went through all 74 AI-visibility and GEO software tools in our index (the 15 managed citation-services agencies sell a different thing, a team rather than a plan, so they’re excluded here, consistent with how this site’s other census articles scope the same corpus) and checked whether each vendor publishes a genuine, checkable methodology for its own accuracy, or just asserts one.

6 of 74 (8%) publish a real, checkable methodology: a sample size, a formula, an open dataset, or a published study a reader could actually verify. 36 of 74 (49%) make an accuracy or precision claim with no evidence behind it. 32 of 74 (43%) make no accuracy claim at all, positive or negative.

How we classified each tool

Every listing in our index is researched and sourced against the vendor’s own live pages. For this census we went back to each tool’s own site (not a new vendor survey, a fresh methodology-specific check against sourced and current material) and looked for one specific thing: does the vendor show HOW it knows its own numbers are right, not just that a customer got results. Three buckets:

  1. Public methodology. A dedicated page, paper, dataset or open-source method disclosing a sample size, a formula, or something a skeptical reader could reproduce or audit.
  2. Claims only. The marketing copy asserts accuracy, precision or “real” data (a specific confidence interval, a scale figure, a “we do it right” pitch) with nothing behind it a reader can check.
  3. Not found. No accuracy or precision claim of any kind, positive or negative. This isn’t a strike against the tool. It’s the honest, largest bucket: silence, not a finding.

The 6 that show their work

Even inside this small group, “shows its work” means very different things. Two disclose their own accuracy numbers directly; one references independent outside research; one publishes an open dataset and paper; one is a published study justifying a design choice rather than a validated accuracy rate; one is code, not prose.

Tool Category What it publishes
EdenRank Visibility/GEO optim. EdenRank’s /proof page is a live, auto-updating counter: 559 of 6,846 answers it measured across 10 engines over the trailing 60 days cited edenrank.com (8%), with a month-by-month log and specific cited-prompt examples anyone can check today.
Evertune Visibility/GEO optim. Evertune’s own methodology page discloses sampling each prompt 100 times across every model, reports how margin of error falls from about ±9 points on a single sample to about ±1 point at 100x, and names “EverPanel,” a 150M-conversation weighted panel behind its brand-perception numbers.
RankLens Visibility/Rank tracking RankLens backs its scoring with a DOI-registered, open-access paper (15,600 samples across 52 categories and 4 locales), publishing its “Entity-Conditioned Probing” method, an overlap@k stability metric, and the underlying code and dataset for anyone to rerun.
RankScale Visibility/GEO optim./Rank tracking RankScale’s /facts page names its “Prompt Decoding” method (credited to a named researcher) and anchors it to an independently checkable, externally published source (an NBER working paper on how people use ChatGPT) rather than an internal number nobody else can see.
Surfer SEO GEO optim./Visibility Surfer SEO published its own 1,000-prompt study comparing API-based and browser-scraped ChatGPT/Perplexity answers, with named authors and concrete numbers (24% brand overlap, 4-8% source overlap between the two methods) used to justify why its own tracker scrapes the UI.
auto-geo GEO optim./Visibility auto-geo is the narrowest case here: an open-source project whose brand-detection logic (case-insensitive domain matching, subdomain handling, false-positive rejection) is fully readable in its own repo rather than described in prose. No accuracy percentage is published, but the method itself is the code, and anyone can audit it directly.

The 36 that claim it and don’t show it

This is the largest and most interesting group: vendors who chose to make an accuracy or precision claim, then didn’t back it with anything a reader could check. A recurring shape shows up across more than a third of this group: a specific-sounding number (“95% confidence interval,” “99.9% accuracy,” “110,504 audits”) that turns out to describe something adjacent to the claim (a different product line, a customer testimonial, a scale metric) rather than a validated accuracy rate for the AI-visibility feature itself.

Tool Category The claim, and what’s missing
AI Sightline Visibility/Rank tracking Cites a third-party statistic about competitors missing mobile clicks, but discloses no methodology or accuracy evidence for its own composite visibility score.
AI Visibility Report Group Visibility Its “Meetmethode” section and knowledge base argue that a standardised pattern of question types delivers “minder bias, meer vergelijkbaarheid” (less bias, more comparability) than a single prompt, and it names its prompt categories (unbranded, category, comparison, problem and purchase-intent questions). The scenario counts it publishes (±25 and ±50) are the size of the report you buy at each tier, not a validation sample, and no formula is disclosed for the readiness score it outputs.
AIclicks Visibility/GEO optim. States it queries AI platforms through their real user interfaces rather than APIs and cites one internal case study, but that’s a customer-outcome anecdote, not a published accuracy check.
Adobe Brand Visibility Visibility/GEO optim. Cites “261 million real AI search prompts (not modeled estimates)” as its data source, but never discloses the scoring formula or a validation figure for its own visibility score.
Ahrefs Brand Radar Visibility Touts “476M+ total monthly prompts” and describes them as “search-backed prompts, not synthetic ones,” but publishes no formula or confidence method (its own /methodology path returns a 404).
Canonry Visibility/GEO optim. Links a dedicated /aeo-methodology page and says it “reads trends across many runs, not single answers,” but the page names no run count, formula, or sample.
CiteLens Visibility/GEO optim. Advertises a “95% confidence interval” on its own scores on the homepage, but its linked /research page contains only unrelated third-party studies, not the sampling behind that number.
Directree GEO Monitor Visibility Claims “every full response is stored as evidence” toward “one visibility score you can explain,” but never discloses the score’s aggregation formula or sample size.
GEO Tool GEO optim./Visibility Has a dedicated /methodik page distinguishing “real queries on live AI interfaces” from competitors’ estimates and cites an external academic paper for credibility, but discloses no sample size or validation figure for its own product.
Gauge Visibility/GEO optim. Claims its data comes from “real data from the true web experiences” and implies rivals get it wrong, but shows no methodology, sample size, or validation for that claim.
Goodie AI GEO optim./Visibility Runs a public “Research Lab” with case-study numbers like a “127% increase in AI conversions,” but nothing in it explains how Goodie itself measures visibility accuracy.
HubSpot AEO Visibility/GEO optim. Describes its mechanism (daily prompts across ChatGPT, Gemini and Perplexity, scored for visibility, citation and sentiment) and a customer testimonial, but no validation data for the scoring itself.
Keyword.com AI Visibility Rank tracking/Visibility The strongest case in this group, and it still lands here for the same reason Nightwatch does. Keyword.com publishes a free 2024 rank-tracker accuracy report with a documented, reproducible test design, and a Spyglass Verification product built on stored HTML SERP snapshots. Both are scoped to its Google rank tracker. Its AI-visibility side is sold as “one accurate AI rank tracker” with no sample size, formula or validation figure of its own.
Knowatoa Visibility/GEO optim. Names a proprietary “BISCUIT Framework,” described as “like PageRank for AI,” and cites “110,504 audits completed,” but discloses no formula or accuracy check behind it.
LLM Pulse Visibility/Rank tracking States it runs “thousands of prompts” weekly with a “28-day rolling aggregate” for normalization, but gives no sample-size breakdown or validation against a known-correct answer.
LLMrefs Visibility/Rank tracking Calls itself “one of the only accurate tools on the market” with “statistically significant” outcomes, but links no methodology page or sample data anywhere.
Local Dominator AI Tracker Visibility/Rank tracking “Most accurate” claims are backed only by unverified customer testimonials, with no methodology or benchmark page.
Local Falcon Rank tracking/Visibility Has a dedicated “Is Local Falcon Accurate?” page, the closest thing to a real attempt in this group, but it’s a two-example anecdotal comparison against one rival, not a sample or formula.
Mangools AI Search Watcher Rank tracking/Visibility Says “running each prompt multiple times ensures accurate averages” and claims “accurate, repeatable insights,” but publishes no methodology page or verifiable data behind either claim.
Mentionable Visibility/GEO optim. Its homepage states scans are “built on the real answers from ChatGPT, Gemini and Perplexity. Not a made-up score,” an implicit real-vs-synthetic accuracy claim, but discloses no sample size, formula or methodology page behind it.
Meltwater GenAI Lens Visibility Advertises “real-time 360° visibility” with product tours and testimonials, but no methodology or accuracy documentation.
Morningscore Visibility/Rank tracking/GEO optim. Describes its GEO Score as tracking “100 prompts,” a feature spec rather than an accuracy proof, and touts “high quality data” with no formula or validation study behind it.
Nightwatch Rank tracking/Visibility Advertises “99.9% accuracy,” but that number is scoped to its legacy SERP rank-tracking product (backed by raw HTML snapshots), not its AI-visibility/LLM-citation feature, which carries no equivalent figure.
Otterly AI Visibility/Rank tracking Its FAQ calls itself “the most neutral, objective monitoring available” and cites scale (“millions of AI citations daily”), but no sample size, formula, or dedicated accuracy page exists.
Promptwatch Visibility/GEO optim. Says it collects data by “scraping the UI interfaces of the LLMs” and cites a large aggregate figure (4.5 billion+ citations, clicks and prompts) to imply scale, but discloses no sample size or validation method.
Qwairy Visibility/GEO optim. The homepage literally promises “clear methodology, precise data” and says “you see exactly how we track, what we measure,” then never actually shows it: no sample size, formula, or checkable page exists.
RadarKit Visibility/GEO optim. Claims greater accuracy from using “real browser sessions” instead of APIs across 40-50+ countries, but that is a feature pitch, not a published methodology with numbers.
Rank Prompt Visibility/GEO optim. Describes a “real-scan” browser-capture approach and an “AI Visibility Score,” but discloses no sample size or validation data on its own site (a circulating “95%+ accuracy” figure lives only on a third-party review site, not on rankprompt.com).
SEORCE Visibility/Rank tracking/GEO optim. Claims “279M+ AI prompts tracked” and an “80% avg. lift,” but its own on-page “See the methodology” link returns a 404; the closest live content discusses AI-answer volatility generally, not SEORCE’s own sample or accuracy.
Scrunch AI Visibility Its FAQ describes a browser-automation-plus-API collection process in prose and claims validation “against a large, continuously updated dataset,” but gives no sample size or formula; a companion blog post on volume estimates explicitly withholds its formula, calling the output “a compass, not a GPS.”
Semrush AI Visibility Toolkit Rank tracking/GEO optim. Its own knowledge-base article discloses real scale (289M+ prompts, 40+ regional databases) and the scoring inputs (topic coverage times mention frequency) but explicitly states “no platform can provide exact numbers on visibility”: real transparency about the data source, not a validated accuracy figure.
Sleepwalker Visibility/GEO optim. Promises to “measure your brand’s visibility with precision” but shows no sample size, formula, or validation evidence anywhere on the site.
Vismore Visibility/GEO optim. Publishes a customer testimonial calling its recommendations “incredibly accurate” alongside headline result stats like a “78% AI Answer Visibility Lift,” with no methodology or sample size behind either.
Webglazer Visibility Its FAQ asks “How do you measure without making the numbers up?” and answers that it “re-run[s] your prompts over time and report[s] appearance rate and share of voice, not one lucky result” — an implicit reliability claim, the same shape as Mentionable’s “not a made-up score” above — but discloses no sample size, formula or dedicated methodology page behind it.
Writesonic Visibility/GEO optim. Claims a “2 billion+ real AI conversations” dataset across 10 platforms and 50+ markets via an “ensemble approach,” but discloses no validation, accuracy metric, or source list for that dataset.
ZipTie Visibility/GEO optim. Argues browser-based capture beats APIs and says it “prioritizes accuracy over convenience,” but supports that only by citing other companies’ published research, not any accuracy check of ZipTie itself.

The 32 that make no claim either way

Addlly AI, CheckThat.ai, TrueRanker, Rankfender, Amplitude AI Visibility, Analyze AI, Archytas AISpy, AthenaHQ, BabyLoveGrowth, Big Leads, Bloomiro, Brandlight, Cision AI Visibility Dashboard, CiteMentor, Cognizo, Conductor, Dageno AI, Frase, GEOrank, GeoRankers, Omnia, Peec AI, Profound, Rankshift, SE Visible, Searchable, Similarweb AI Search Intelligence, Syntropic AI Visibility Audit, Trakkr, Ubersuggest, VisibAI, geoSurge simply don’t raise accuracy or precision as a topic on their own sites, one way or the other. That’s not evidence of anything: a tool can be accurate and never say so, the same way a tool can claim 99% accuracy and be wrong. Absence of a claim just means this census has nothing to report for it.

Why this is a real buying question

  • A specific number is not evidence. “95% confidence interval,” “110,504 audits,” “279M+ prompts tracked”: these read like methodology, and several vendors clearly intend them to. None of them, on their own, tell you how the vendor knows its citation count is right. Ask what the number would look like if the tool were wrong, and whether the vendor has ever shown that check.
  • The strongest signal is a page you could argue with. EdenRank’s /proof page and Evertune’s methodology page both let you find a specific number and ask “is that enough of a sample.” RankLens goes further and publishes the underlying dataset. That’s a different category of claim than a homepage adjective, and it’s rare among the 6 tools in this census that publish anything at all.
  • A missing claim isn’t a red flag by itself. Don’t read the 32-tool “not found” group as 32 tools hiding something. Plenty of legitimate products simply don’t lead with an accuracy pitch. Weigh it alongside everything else in the listing, not as a standalone strike.

What this census doesn’t tell you

This counts whether a vendor publishes checkable evidence for its own accuracy claim, not whether the tool is actually accurate. A vendor with no public methodology page could still be measuring correctly; a vendor with a polished-looking methodology page could still be wrong in ways this census can’t detect from the outside, since we didn’t independently re-run anyone’s measurement. It’s also a snapshot: sites change, and a “claims only” or “not found” classification today could be a “public methodology” one after a future product update. If a vendor’s measurement accuracy is a deciding factor in your purchase, ask them directly for the evidence this census didn’t find on their site, and check whether their answer is something you could actually verify.

Methodology

We visited each of the 74 tools’ own live sites (homepage plus, where it existed, a linked methodology, proof, research or FAQ page) and searched specifically for a disclosed sample size, formula, dataset, published study or open-source method behind the vendor’s own accuracy or precision claims, as distinct from product-feature descriptions, customer testimonials, or scale claims about data volume alone. A claim found only on a third-party review site, not the vendor’s own domain, does not count as the vendor showing its work, and is noted separately where it changed a classification. The 74-tool base matches this index’s tools-only corpus on 16 August 2026 (89 published listings, less the 15 citation-services agencies), the same corpus this site’s engine-coverage and free-tier censuses draw from, so results can be cross-referenced against either. Note that those two report out of a smaller denominator than this one: engine-coverage answers its question for 72 of the 74, setting Big Leads and AI Visibility Report Group aside as undetermined, and the free-tier census excludes Big Leads as a growth agency and is still reported against an earlier base that this pass has not rebased. Every tool in this census could be classified, because “does this vendor publish a methodology?” has an answer for a service business as much as for software, so the base here is the full 74. This is a live-site snapshot rather than a continuously monitored feed: as with our other census articles, a future pass will note when a tool’s classification changes rather than silently overwrite it.

2026-08-16 update (ninth pass): rebased from 72 to the live 74-tool corpus. The two listings published since the eighth pass were checked against their own live sites today and both join “claims only”, which moves from 34 to 36; the 6-tool “public methodology” and 32-tool “not found” groups are unchanged, giving 6 + 36 + 32 = 74. Shares are restated against the new base (8% / 49% / 43%); no tool already in this census changed bucket. Keyword.com AI Visibility is the more interesting of the two and the strongest “claims only” case in the article: it genuinely publishes a free 2024 rank-tracker accuracy report with a documented, reproducible test design, and sells a third-party SERP verification product backed by stored HTML snapshots. Both are scoped to its Google rank tracker, while its AI-visibility side is sold as “one accurate AI rank tracker” with no sample size, formula or validation of its own — the same split that keeps Nightwatch’s “99.9% accuracy” out of the top group. AI Visibility Report Group publishes a “Meetmethode” section and a knowledge base arguing that a standardised pattern of question types gives less bias and more comparability than a single prompt, and names its prompt categories, but the scenario counts it discloses (±25 and ±50) size the report you buy rather than validate the score, and no formula for that score is published — the same distinction that keeps Morningscore’s “100 prompts” and LLM Pulse’s “thousands of prompts” in this group. Every name in all three buckets was set-diffed against the live corpus this pass: all 72 previously listed tools are still published, none is duplicated, and these two were the only unclassified ones.

2026-08-13 update (eighth pass): rebased from 70 to the live 72-tool corpus. Two listings published since the seventh pass had never been classified here: Big Leads (a growth agency carrying only the geo-optimization category, so it counts inside the tools-only base) and CiteMentor. Both were checked today against their own live sites, bigleads.io home, /about, /services and /case-studies, plus citementor.ai home, /about, /platform and /pricing, all returning HTTP 200. Neither raises accuracy or precision as a topic anywhere: no sample size, no formula, no validation figure, and equally no “most accurate” pitch to hold against them. Both join “not found”, which moves from 30 to 32; the 6-tool “public methodology” and 34-tool “claims only” groups are unchanged, giving 6 + 34 + 32 = 72. Shares are restated against the new base (8% / 47% / 44%). That is the only reason those figures moved; no tool changed bucket. The method paragraph now also spells out where this census’s denominator differs from the engine-coverage and free-tier censuses, which set aside one and two of these same 72 respectively; that divergence was previously implied to be nil.

_2026-08-11 update (seventh pass): resolved the standing gap the sixth pass could only name. Two separate bugs, not one: (1) Webglazer had been counted into corpusCountTools since 2026-08-10/11 but never classified — its own FAQ (“How do you measure without making the numbers up?”) makes the same implicit-reliability claim as Mentionable’s “not a made-up score” with no sample size or formula behind it, so it joins the “claims only” group; “claims only” moves from 33 to 34. (2) Cross-checking the full 70-tool live corpus against every name in all three buckets turned up a second, older bug: Brandlight was announced as added to “not found” in the 2026-08-09 (second) pass, but never actually appended to the printed list — a write that described itself without happening. Added now; “not found” stays at 30 (it already silently carried Brandlight’s slot, unnamed) but is now honestly enumerable: 6 + 34 + 30 = 70, matching corpusCountTools for the first time since this census began carrying a phantom count. Neither correction is a re-research of either tool’s site; both rest on evidence already quoted elsewhere in this article or, for Webglazer, a fresh FAQ check against its live site today.

_2026-08-11 update (sixth pass): added Rankfender, an AI-visibility-plus-content-generation platform from agency 361 DEV, to the “not found” group — its home, pricing and features pages describe RAIVE’s scoring (visibility score, share of voice, citation rate) in detail but make no accuracy or precision claim, positive or negative, about how well that scoring matches reality. Base moves from 69 to 70 tools; the 6-tool “public methodology” and 33-tool “claims only” groups are both unchanged; “not found” moves from 29 to 30. This pass also confirmed the pre-existing gap noted below the list is Webglazer, not a mystery name: it was counted into this article’s corpusCountTools on 2026-08-11 alongside the other 6 corpus-count-gated articles but never actually researched for this specific census — left unresolved rather than guessed at, per the same standing debt the 2026-08-10 pass flagged.

2026-08-10 update (fifth pass): added TrueRanker, a bootstrapped SEO rank tracker (est. 2019) with an AI-visibility module, to the “not found” group — its pricing and product pages describe keyword/AI-prompt caps and features in granular detail but make no accuracy or precision claim, positive or negative, about its own AI-visibility measurement. Base moves from 67 to 68 tools; the 6-tool “public methodology” and 33-tool “claims only” groups are both unchanged; “not found” moves from 28 to 29 (see the note above this list about a pre-existing 1-name enumeration gap found during this pass, not introduced by it). 2026-08-10 update (fourth pass): added CheckThat.ai, a free AI-visibility benchmarking platform built by GrowthX, to the “not found” group — its homepage emphasizes scale (2.6M+ tracked AI responses, 5,983 brands) and openness, but makes no accuracy or precision claim, positive or negative, about how well its own tracking matches reality; scale-of-data claims alone don’t count per this census’s own methodology, the same distinction applied to Adobe Brand Visibility and Ahrefs Brand Radar above. Base moves from 66 to 67 tools; the 6-tool “public methodology” and 33-tool “claims only” groups are both unchanged; “not found” moves from 27 to 28. 2026-08-09 update (third pass): added Cognizo, an AEO platform bundling AI-visibility monitoring with content production, to the “not found” group — its homepage schema markup describes running prompts continuously “to collect millions of responses daily” but makes no accuracy, precision or real-vs-estimated claim about its own scores anywhere checked. Base moves from 65 to 66 tools; the 6-tool “public methodology” and 33-tool “claims only” groups are both unchanged; “not found” moves from 26 to 27. 2026-08-09 update (second pass): added Brandlight, a well-funded enterprise AI-visibility platform, to the “not found” group — its homepage leans on funding and named-customer credibility (“the obvious enterprise choice,” CB Insights/Gartner recognition) rather than any accuracy or precision claim about its own measurement, so it makes no claim to evaluate either way. Base moves from 64 to 65 tools; the 6-tool “public methodology” and 33-tool “claims only” groups are both unchanged; “not found” moves from 25 to 26. 2026-08-09 update: added Mentionable, a credit-tiered AI-visibility monitor, to the “claims only” group — its homepage’s “not a made-up score” line is an implicit accuracy claim (real vs. synthetic data) with no disclosed sample size or formula behind it, the same shape as Adobe Brand Visibility’s “not modeled estimates” framing above. Base moves from 63 to 64 tools; the “public methodology” group stays at 6, “claims only” moves from 32 to 33, “not found” stays at 25. 2026-08-08 update: added Omnia, an agentic AI-visibility platform, to the “not found” group — its pricing and product pages describe its real, geo-located browser-check methodology in detail (specific countries and languages, a 24-hour refresh cycle with a moving average) but stop short of an explicit accuracy or precision claim, positive or negative, the same distinction that keeps RadarKit and ZipTie in the “claims only” group above rather than this one; if a future check finds Omnia explicitly asserting greater accuracy from that approach, it moves up to that group instead. Base moves from 62 to 63 tools; the 6-tool “public methodology” and 32-tool “claims only” groups are both unchanged. 2026-08-07 update: added Analyze AI, a GA4-attribution-focused GEO platform, to the “not found” group — its homepage, pricing and Discover/Monitor/Improve/Govern feature pages describe its tracking and content-optimization workflow in detail but make no accuracy or precision claim, positive or negative, about its own measurement. Base moves from 61 to 62 tools; the 6-tool “public methodology” and 32-tool “claims only” groups are both unchanged.

Get the next report

New tools rankings and fresh data reports. One short email, straight to your inbox. One-click unsubscribe.