Research report · geo
Which AI-Visibility Tools Show Their Work? A Methodology Census
A dated 74-tool transparency review: six had inspectable supporting material, not independently validated accuracy. Corrected evidence scope and sampling units.
On this page
AI-visibility scores are buying decisions only if their measurement is understandable. This is the retained 16 August 2026 editorial review of 74 named tools, excluding the 15 entries then classified as managed citation services. Its useful question is what supporting material the review found on public pages, not which tool has proved an accurate score. The six positive examples below have been source-checked for this correction; the 74-tool classification exercise has not been re-run, and we do not assert that all 74 remain published or retain the same taxonomy today.
The historical classification was 6 of 74 (8%) with public methods or supporting material, 36 of 74 (49%) with claims but no qualifying support found in the reviewed pages, and 32 of 74 (43%) with no qualifying accuracy claim found. These are editorial disclosure buckets, not measured accuracy rates or proof that evidence did not exist elsewhere.
How we classified each tool
The original review describes looking at homepages and linked methodology, proof, research or FAQ pages for sample sizes, formulas, datasets, studies or code. It did not independently execute the tools or test their scores against a ground truth. The exact contemporaneous page captures and a reproducible row-level classification ledger were not recovered for all 74 names; immutable article revision 22 preserves the published classifications, not the underlying vendor evidence. Read the negative findings as bounded observations from that review, not an exhaustive search of every page on each domain.
- Public methods or supporting material. A page, paper, dataset or code that gives a reader something to inspect. This broad bucket includes both sampling methodology and much weaker material such as a citation counter or an outside usage study; those are not equivalent validation.
- Claims only in the reviewed material. An accuracy, precision or real-data pitch for which the review did not find qualifying support on the pages it examined. Customer outcomes and large volume counters do not alone validate measurement accuracy.
- Claim not found in the reviewed material. The review did not identify an accuracy or precision claim to assess. This is not evidence against the product. The 36-tool claims-only group, not this 32-tool group, was the largest.
The six recorded as having inspectable supporting material
The distinction inside this group matters more than membership. Evertune describes a sampling procedure; RankLens publishes a research project; Surfer compares two collection surfaces; auto-geo documents inspectable code. EdenRank’s own-brand citation frequency is an outcome count, not measurement accuracy, and RankScale’s cited outside study is about ChatGPT usage, not validation of RankScale. None becomes independent product validation merely by being linked here.
| Tool | Category | What it publishes |
|---|---|---|
| EdenRank | Visibility/GEO optim. | The edition recorded 559/6,846 answers (about 8%) citing edenrank.com across 10 engines in a trailing 60-day counter on its /proof page. The counter is mutable and the original window endpoints were not retained here; today’s total cannot verify that historical rate. The page’s log and examples expose claimed citation occurrences, not a false-positive/false-negative accuracy test. |
| Evertune | Visibility/GEO optim. | Its methodology page gives a July 2026 running-shoe example: 100 distinct prompts asked once (100 responses) produced an overall margin of about ±9 points; repeating each of those 100 prompts 100 times (10,000 responses) produced about ±1 overall. This is not ±1 for one prompt repeated 100 times: the page separately illustrates about ±10 for that case. EverPanel’s claimed 150M-conversation weighting is a vendor account, not independently confirmed representativeness. |
| RankLens | Visibility/Rank tracking | Its OSF project, with public metadata, describes 15,600 prompt iterations across 52 categories and four locales, Entity-Conditioned Probing, resampling and stability checks. The project description says code and processed aggregates are open; a DOI does not imply peer review, complete raw responses, or that this review reproduced the product’s accuracy. |
| RankScale | Visibility/GEO optim./Rank tracking | Its /facts page names Hanns Kronenberg’s Prompt Decoding method and an exclusive licence. It compares its findings with NBER working paper 34255, How People Use ChatGPT, a study of ChatGPT conversation topics. That paper does not test RankScale’s visibility scores; agreement claimed by RankScale is not independent validation of RankScale. |
| Surfer SEO | GEO optim./Visibility | Its API-versus-browser study describes 1,000 executions per scenario: 24% ChatGPT brand overlap, 4% ChatGPT source overlap and 8% Perplexity source overlap. These are vendor-reported overlaps for those scenarios, not universal accuracy rates. The study supports its own browser-based tracker and was not independently replicated here. |
| auto-geo | GEO optim./Visibility | Its check documentation describes engine adapters and brand/domain matching in an open-source implementation. Inspectability lets a reader examine the rules; it does not demonstrate that every reported citation is correct or that API output matches the consumer interface. No execution or accuracy replication was performed in this correction. |
The 36 historical claims-only classifications
The following rows preserve the substantive negative findings of the 16 August review. A score’s volume counter, a customer success story and a validation dataset answer different questions. Each row reports what that review said was missing from its inspected material; it is not a present-day certification that no documentation exists. The retained internal listing links are navigational references, not immutable copies of the old vendor pages. The original review’s failure to retain all page-level evidence limits independent verification of these classifications.
| Tool | Category | Recorded 16 August finding; not a current documentation audit |
|---|---|---|
| AI Sightline | Visibility/Rank tracking | Cites a third-party statistic about competitors missing mobile clicks, but discloses no methodology or accuracy evidence for its own composite visibility score. |
| AI Visibility Report Group | Visibility | Its “Meetmethode” section and knowledge base argue that a standardised pattern of question types delivers “minder bias, meer vergelijkbaarheid” (less bias, more comparability) than a single prompt, and it names its prompt categories (unbranded, category, comparison, problem and purchase-intent questions). The scenario counts it publishes (±25 and ±50) are the size of the report you buy at each tier, not a validation sample, and no formula is disclosed for the readiness score it outputs. |
| AIclicks | Visibility/GEO optim. | States it queries AI platforms through their real user interfaces rather than APIs and cites one internal case study, but that’s a customer-outcome anecdote, not a published accuracy check. |
| Adobe Brand Visibility | Visibility/GEO optim. | Cites “261 million real AI search prompts (not modeled estimates)” as its data source, but the review found no disclosure of the scoring formula or a validation figure for its own visibility score. |
| Ahrefs Brand Radar | Visibility | Touts “476M+ total monthly prompts” and describes them as “search-backed prompts, not synthetic ones,” but the review recorded no formula or confidence method and a 404 at /methodology (not a current HTTP check). |
| Canonry | Visibility/GEO optim. | Links a dedicated /aeo-methodology page and says it “reads trends across many runs, not single answers,” but the page names no run count, formula, or sample. |
| CiteLens | Visibility/GEO optim. | Advertises a “95% confidence interval” on its own scores on the homepage, but its linked /research page contains only unrelated third-party studies, not the sampling behind that number. |
| Directree GEO Monitor | Visibility | Claims “every full response is stored as evidence” toward “one visibility score you can explain,” but the review found no disclosure of the score’s aggregation formula or sample size. |
| GEO Tool | GEO optim./Visibility | Has a dedicated /methodik page distinguishing “real queries on live AI interfaces” from competitors’ estimates and cites an external academic paper for credibility, but discloses no sample size or validation figure for its own product. |
| Gauge | Visibility/GEO optim. | Claims its data comes from “real data from the true web experiences” and implies rivals get it wrong, but shows no methodology, sample size, or validation for that claim. |
| Goodie AI | GEO optim./Visibility | Runs a public “Research Lab” with case-study numbers like a “127% increase in AI conversions,” but nothing in it explains how Goodie itself measures visibility accuracy. |
| HubSpot AEO | Visibility/GEO optim. | Describes its mechanism (daily prompts across ChatGPT, Gemini and Perplexity, scored for visibility, citation and sentiment) and a customer testimonial, but no validation data for the scoring itself. |
| Keyword.com AI Visibility | Rank tracking/Visibility | The strongest case in this group, and it still lands here for the same reason Nightwatch does. Keyword.com publishes a free 2024 rank-tracker accuracy report with a documented, reproducible test design, and a Spyglass Verification product built on stored HTML SERP snapshots. Both are scoped to its Google rank tracker. Its AI-visibility side is sold as “one accurate AI rank tracker” with no sample size, formula or validation figure of its own. |
| Knowatoa | Visibility/GEO optim. | Names a proprietary “BISCUIT Framework,” described as “like PageRank for AI,” and cites “110,504 audits completed,” but discloses no formula or accuracy check behind it. |
| LLM Pulse | Visibility/Rank tracking | States it runs “thousands of prompts” weekly with a “28-day rolling aggregate” for normalization, but gives no sample-size breakdown or validation against a known-correct answer. |
| LLMrefs | Visibility/Rank tracking | Calls itself “one of the only accurate tools on the market” with “statistically significant” outcomes, but the review found no linked methodology page or sample data in its inspected material. |
| Local Dominator AI Tracker | Visibility/Rank tracking | “Most accurate” claims are backed only by unverified customer testimonials, with no methodology or benchmark page. |
| Local Falcon | Rank tracking/Visibility | Has a dedicated “Is Local Falcon Accurate?” page, the closest thing to a real attempt in this group, but it’s a two-example anecdotal comparison against one rival, not a sample or formula. |
| Mangools AI Search Watcher | Rank tracking/Visibility | Says “running each prompt multiple times ensures accurate averages” and claims “accurate, repeatable insights,” but publishes no methodology page or verifiable data behind either claim. |
| Mentionable | Visibility/GEO optim. | Its homepage states scans are “built on the real answers from ChatGPT, Gemini and Perplexity. Not a made-up score,” an implicit real-vs-synthetic accuracy claim, but discloses no sample size, formula or methodology page behind it. |
| Meltwater GenAI Lens | Visibility | Advertises “real-time 360° visibility” with product tours and testimonials, but no methodology or accuracy documentation. |
| Morningscore | Visibility/Rank tracking/GEO optim. | Describes its GEO Score as tracking “100 prompts,” a feature spec rather than an accuracy proof, and touts “high quality data” with no formula or validation study behind it. |
| Nightwatch | Rank tracking/Visibility | Advertises “99.9% accuracy,” but that number is scoped to its legacy SERP rank-tracking product (backed by raw HTML snapshots), not its AI-visibility/LLM-citation feature, which carries no equivalent figure. |
| Otterly AI | Visibility/Rank tracking | Its FAQ calls itself “the most neutral, objective monitoring available” and cites scale (“millions of AI citations daily”), but no sample size, formula, or dedicated accuracy page was found in the reviewed material. |
| Promptwatch | Visibility/GEO optim. | Says it collects data by “scraping the UI interfaces of the LLMs” and cites a large aggregate figure (4.5 billion+ citations, clicks and prompts) to imply scale, but discloses no sample size or validation method. |
| Qwairy | Visibility/GEO optim. | The homepage literally promises “clear methodology, precise data” and says “you see exactly how we track, what we measure,” but the review found no sample size, formula, or checkable supporting page in its inspected material. |
| RadarKit | Visibility/GEO optim. | Claims greater accuracy from using “real browser sessions” instead of APIs across 40-50+ countries, but that is a feature pitch, not a published methodology with numbers. |
| Rank Prompt | Visibility/GEO optim. | Describes a “real-scan” browser-capture approach and an “AI Visibility Score,” but discloses no sample size or validation data on its own site (a circulating “95%+ accuracy” figure lives only on a third-party review site, not on rankprompt.com). |
| SEORCE | Visibility/Rank tracking/GEO optim. | Claims “279M+ AI prompts tracked” and an “80% avg. lift,” but the review recorded a 404 from its on-page “See the methodology” link; the closest live content discusses AI-answer volatility generally, not SEORCE’s own sample or accuracy. |
| Scrunch AI | Visibility | Its FAQ describes a browser-automation-plus-API collection process in prose and claims validation “against a large, continuously updated dataset,” but gives no sample size or formula; a companion blog post on volume estimates explicitly withholds its formula, calling the output “a compass, not a GPS.” |
| Semrush AI Visibility Toolkit | Rank tracking/GEO optim. | Its own knowledge-base article discloses real scale (289M+ prompts, 40+ regional databases) and the scoring inputs (topic coverage times mention frequency) but explicitly states “no platform can provide exact numbers on visibility”: real transparency about the data source, not a validated accuracy figure. |
| Sleepwalker | Visibility/GEO optim. | Promises to “measure your brand’s visibility with precision” but shows no sample size, formula, or validation evidence in the material inspected by that review. |
| Vismore | Visibility/GEO optim. | Publishes a customer testimonial calling its recommendations “incredibly accurate” alongside headline result stats like a “78% AI Answer Visibility Lift,” with no methodology or sample size behind either. |
| Webglazer | Visibility | Its FAQ asks “How do you measure without making the numbers up?” and answers that it “re-run[s] your prompts over time and report[s] appearance rate and share of voice, not one lucky result” — an implicit reliability claim, the same shape as Mentionable’s “not a made-up score” above — but discloses no sample size, formula or dedicated methodology page behind it. |
| Writesonic | Visibility/GEO optim. | Claims a “2 billion+ real AI conversations” dataset across 10 platforms and 50+ markets via an “ensemble approach,” but discloses no validation, accuracy metric, or source list for that dataset. |
| ZipTie | Visibility/GEO optim. | Argues browser-based capture beats APIs and says it “prioritizes accuracy over convenience,” but supports that only by citing other companies’ published research, not any accuracy check of ZipTie itself. |
The 32 for which the review found no qualifying claim
Addlly AI, CheckThat.ai, TrueRanker, Rankfender, Amplitude AI Visibility, Analyze AI, Archytas AISpy, AthenaHQ, BabyLoveGrowth, Big Leads, Bloomiro, Brandlight, Cision AI Visibility Dashboard, CiteMentor, Cognizo, Conductor, Dageno AI, Frase, GEOrank, GeoRankers, Omnia, Peec AI, Profound, Rankshift, SE Visible, Searchable, Similarweb AI Search Intelligence, Syntropic AI Visibility Audit, Trakkr, Ubersuggest, VisibAI, geoSurge were recorded as having no qualifying accuracy or precision claim in the pages inspected by the 16 August review. That’s not evidence of anything: a tool can be accurate and never say so, the same way a tool can claim 99% accuracy and be wrong. Absence of a claim just means this census has nothing to report for it.
Why this is a real buying question
- A specific number is not validation by itself. “95% confidence interval,” “110,504 audits,” “279M+ prompts tracked” describe different things. Ask for the unit, sampling process, error calculation and a check against a defined reference; do not infer intent or accuracy from a headline number.
- Inspectable material is a starting point, not a pass mark. EdenRank’s counter, Evertune’s sampling example and RankLens’s research project expose different evidence. A citation frequency cannot substitute for a scoring-error measurement, and processed aggregates are not necessarily a complete replication dataset.
- A missing claim isn’t a red flag by itself. Don’t read the 32-tool “not found” group as 32 tools hiding something. Plenty of legitimate products simply don’t lead with an accuracy pitch. Weigh it alongside everything else in the listing, not as a standalone strike.
What this census doesn’t tell you
This article preserves a historical transparency inventory, not an accuracy ranking. The top bucket mixes evidence of different strengths and should not be interpreted as six independently validated products versus 68 untrustworthy ones. A missing public page does not prove inaccurate measurement; a polished methodology does not prove correctness. Sites can change, and the record here does not establish when later documentation appeared. Buyers should inspect the current methods and ask whether sampling, scoring and raw evidence support the specific decision they intend to make.
Methodology
The recorded census base was 74 tools on 16 August 2026, described by the original edition as 89 published listings less 15 citation-services entries. Different companion censuses excluded additional ambiguous entries and were updated on different days; their percentages must not be treated as one synchronized cohort. Revision 22, written 16 August at 06:01 UTC, preserves this article’s 6/36/32 classifications. The later warehouse run on that date, 16589 at 21:59 UTC, already contains 91 published entries and 75 non-citation-services listings, so it cannot establish the earlier classification population merely because the dates match. We have not replaced the historical 74 denominator with that later snapshot or today’s directory. The original change log below is retained as the prior editors’ account, including its stated visits and set differences; those actions were not repeated or independently established by this correction.
2026-08-16 update (ninth pass): rebased from 72 to the live 74-tool corpus. The two listings published since the eighth pass were checked against their own live sites today and both join “claims only”, which moves from 34 to 36; the 6-tool “public methodology” and 32-tool “not found” groups are unchanged, giving 6 + 36 + 32 = 74. Shares are restated against the new base (8% / 49% / 43%); no tool already in this census changed bucket. Keyword.com AI Visibility is the more interesting of the two and the strongest “claims only” case in the article: it genuinely publishes a free 2024 rank-tracker accuracy report with a documented, reproducible test design, and sells a third-party SERP verification product backed by stored HTML snapshots. Both are scoped to its Google rank tracker, while its AI-visibility side is sold as “one accurate AI rank tracker” with no sample size, formula or validation of its own — the same split that keeps Nightwatch’s “99.9% accuracy” out of the top group. AI Visibility Report Group publishes a “Meetmethode” section and a knowledge base arguing that a standardised pattern of question types gives less bias and more comparability than a single prompt, and names its prompt categories, but the scenario counts it discloses (±25 and ±50) size the report you buy rather than validate the score, and no formula for that score is published — the same distinction that keeps Morningscore’s “100 prompts” and LLM Pulse’s “thousands of prompts” in this group. Every name in all three buckets was set-diffed against the live corpus this pass: all 72 previously listed tools are still published, none is duplicated, and these two were the only unclassified ones.
2026-08-13 update (eighth pass): rebased from 70 to the live 72-tool corpus. Two listings published since the seventh pass had never been classified here: Big Leads (a growth agency carrying only the geo-optimization category, so it counts inside the tools-only base) and CiteMentor. Both were checked today against their own live sites, bigleads.io home, /about, /services and /case-studies, plus citementor.ai home, /about, /platform and /pricing, all returning HTTP 200. Neither raises accuracy or precision as a topic anywhere: no sample size, no formula, no validation figure, and equally no “most accurate” pitch to hold against them. Both join “not found”, which moves from 30 to 32; the 6-tool “public methodology” and 34-tool “claims only” groups are unchanged, giving 6 + 34 + 32 = 72. Shares are restated against the new base (8% / 47% / 44%). That is the only reason those figures moved; no tool changed bucket. The method paragraph now also spells out where this census’s denominator differs from the engine-coverage and free-tier censuses, which set aside one and two of these same 72 respectively; that divergence was previously implied to be nil.
_2026-08-11 update (seventh pass): resolved the standing gap the sixth pass could only name. Two separate bugs, not one: (1) Webglazer had been counted into corpusCountTools since 2026-08-10/11 but never classified — its own FAQ (“How do you measure without making the numbers up?”) makes the same implicit-reliability claim as Mentionable’s “not a made-up score” with no sample size or formula behind it, so it joins the “claims only” group; “claims only” moves from 33 to 34. (2) Cross-checking the full 70-tool live corpus against every name in all three buckets turned up a second, older bug: Brandlight was announced as added to “not found” in the 2026-08-09 (second) pass, but never actually appended to the printed list — a write that described itself without happening. Added now; “not found” stays at 30 (it already silently carried Brandlight’s slot, unnamed) but is now honestly enumerable: 6 + 34 + 30 = 70, matching corpusCountTools for the first time since this census began carrying a phantom count. Neither correction is a re-research of either tool’s site; both rest on evidence already quoted elsewhere in this article or, for Webglazer, a fresh FAQ check against its live site today.
_2026-08-11 update (sixth pass): added Rankfender, an AI-visibility-plus-content-generation platform from agency 361 DEV, to the “not found” group — its home, pricing and features pages describe RAIVE’s scoring (visibility score, share of voice, citation rate) in detail but make no accuracy or precision claim, positive or negative, about how well that scoring matches reality. Base moves from 69 to 70 tools; the 6-tool “public methodology” and 33-tool “claims only” groups are both unchanged; “not found” moves from 29 to 30. This pass also confirmed the pre-existing gap noted below the list is Webglazer, not a mystery name: it was counted into this article’s corpusCountTools on 2026-08-11 alongside the other 6 corpus-count-gated articles but never actually researched for this specific census — left unresolved rather than guessed at, per the same standing debt the 2026-08-10 pass flagged.
2026-08-10 update (fifth pass): added TrueRanker, a bootstrapped SEO rank tracker (est. 2019) with an AI-visibility module, to the “not found” group — its pricing and product pages describe keyword/AI-prompt caps and features in granular detail but make no accuracy or precision claim, positive or negative, about its own AI-visibility measurement. Base moves from 67 to 68 tools; the 6-tool “public methodology” and 33-tool “claims only” groups are both unchanged; “not found” moves from 28 to 29 (see the note above this list about a pre-existing 1-name enumeration gap found during this pass, not introduced by it). 2026-08-10 update (fourth pass): added CheckThat.ai, a free AI-visibility benchmarking platform built by GrowthX, to the “not found” group — its homepage emphasizes scale (2.6M+ tracked AI responses, 5,983 brands) and openness, but makes no accuracy or precision claim, positive or negative, about how well its own tracking matches reality; scale-of-data claims alone don’t count per this census’s own methodology, the same distinction applied to Adobe Brand Visibility and Ahrefs Brand Radar above. Base moves from 66 to 67 tools; the 6-tool “public methodology” and 33-tool “claims only” groups are both unchanged; “not found” moves from 27 to 28. 2026-08-09 update (third pass): added Cognizo, an AEO platform bundling AI-visibility monitoring with content production, to the “not found” group — its homepage schema markup describes running prompts continuously “to collect millions of responses daily” but makes no accuracy, precision or real-vs-estimated claim about its own scores anywhere checked. Base moves from 65 to 66 tools; the 6-tool “public methodology” and 33-tool “claims only” groups are both unchanged; “not found” moves from 26 to 27. 2026-08-09 update (second pass): added Brandlight, a well-funded enterprise AI-visibility platform, to the “not found” group — its homepage leans on funding and named-customer credibility (“the obvious enterprise choice,” CB Insights/Gartner recognition) rather than any accuracy or precision claim about its own measurement, so it makes no claim to evaluate either way. Base moves from 64 to 65 tools; the 6-tool “public methodology” and 33-tool “claims only” groups are both unchanged; “not found” moves from 25 to 26. 2026-08-09 update: added Mentionable, a credit-tiered AI-visibility monitor, to the “claims only” group — its homepage’s “not a made-up score” line is an implicit accuracy claim (real vs. synthetic data) with no disclosed sample size or formula behind it, the same shape as Adobe Brand Visibility’s “not modeled estimates” framing above. Base moves from 63 to 64 tools; the “public methodology” group stays at 6, “claims only” moves from 32 to 33, “not found” stays at 25. 2026-08-08 update: added Omnia, an agentic AI-visibility platform, to the “not found” group — its pricing and product pages describe its real, geo-located browser-check methodology in detail (specific countries and languages, a 24-hour refresh cycle with a moving average) but stop short of an explicit accuracy or precision claim, positive or negative, the same distinction that keeps RadarKit and ZipTie in the “claims only” group above rather than this one; if a future check finds Omnia explicitly asserting greater accuracy from that approach, it moves up to that group instead. Base moves from 62 to 63 tools; the 6-tool “public methodology” and 32-tool “claims only” groups are both unchanged. 2026-08-07 update: added Analyze AI, a GA4-attribution-focused GEO platform, to the “not found” group — its homepage, pricing and Discover/Monitor/Improve/Govern feature pages describe its tracking and content-optimization workflow in detail but make no accuracy or precision claim, positive or negative, about its own measurement. Base moves from 61 to 62 tools; the 6-tool “public methodology” and 32-tool “claims only” groups are both unchanged.
5 September 2026 editorial correction: corrected Evertune’s aggregation units, separated citation counters and outside usage research from product validation, bounded the negative findings to the inspected material, and corrected which bucket was largest. Original census and publication dates are unchanged. Six positive sources were reviewed; no new 74-tool census, product execution, accuracy experiment or blanket re-verification was performed.
Get the next report
New tools rankings and fresh data reports. One short email, one-click unsubscribe.