New: the nine ways developer tools go invisible in AI answers.
citon

For AI infrastructure, developer tools, APIs, CLIs and MCP servers

Your buyers asked an AI which tool to use. It named your competitor.

We measure where you stand in live AI answers, we show you the exact pages the models cite when your buyers ask, and we go get you onto them. Every engagement runs against a , so at day 90 you get the one number nobody else in this category will give you: how much of the movement we caused.

Live AI answerSample

Your buyer asks

which rate limiting api should we use?

The answer they get

For production workloads, most teams land on Competitor API1. It pairs token-bucket limits with per-key analytics.2

Where the citations resolve

reddit.com38%
g2.com26%
news.ycombinator.com21%
yourapi.exampleNot cited

Measured across

Google AI Overviews
ChatGPT
Claude
Perplexity
Grok
DeepSeek

91.5%

of AI citations point off-site

901 citations across 60 answers. Our own measurement.

82 hosts

carried those citations between them

No single site owns a category, so there is nothing to buy your way onto.

45.5%

of citations change between consecutive observations

Ahrefs, 43,000 keywords. AI Overviews persist 2.15 days.

The AI-citation team for developer tools

We run one play, for one kind of company, all the way down.

Citon works with companies building AI infrastructure, developer tools, APIs, CLIs and MCP servers. Not general B2B SaaS. The surfaces a developer evaluates from are a short and specific list, and it is the same list every time, which is why we work one niche rather than several.

  • One niche
  • Held-out control, every tier
  • A few clients per quarter
  • No dashboard at any tier

A small senior team. No junior handoff after the pitch.

Where developers decide

01

GitHub and Hacker News

Where a tool gets vouched for

02

Docs as answers

Usually the surface that does the most work per hour spent on it

03

Best-in-category roundups

Ranked by topical depth, never by domain rating

04

Community threads

Already being cited, so already worth entering

05

Review sites

The corroboration models trust

Not a marketing blog post

Short answer

How is this different from an AI visibility tool?

A tool reports a score. We give that away free, because monitoring is commoditized and it starts at 99 dollars a month. What you pay for is the work across the surfaces models cite, and a held-out control group showing the work moved the number.

The problem with the current answer

You already know AI decides. Every vendor still sells you a number that moves on its own.

You see a score. Nobody tells you why.

Level measurement is destroyed by non-determinism. We ran twelve queries five times each on one model and seven flipped outcome. That is a property of the system, not a vendor failure.

Your agency is a step behind on AI.

You get a ranking report every month while the clicks behind those rankings fall away. The surfaces that decide an AI answer are mostly not the ones an SEO retainer works on.

Reports tell you what happened. Not what you caused.

Across more than fifty companies mapped in this category, we have not found one publishing a control group or a causal-lift number.

Monitoring is commoditized. It starts at 99 dollars a month, and Google ships AI-search data free inside Search Console. So we give monitoring away and charge for the part nobody else will stand behind.

What changed

Developers stopped browsing results. They ask, and they take the answer.

We run on-site and off-site as one service, because a page with no outside corroboration stays invisible, and corroboration with no page has nothing to point at.

google.com
best rate limiting api
Yyourapi.example › docs↑ 1

Rate limiting for production APIs

competitor.example › blog

roundup.example › guides

Citation share
Q1Q2Q3Q4

On-site

The citable answer

The page the model points at. Necessary, and table stakes. We do this competently and we do not try to out-automate the page engines that already do it at scale.

AAssistant···

Which API should I use for rate limiting?

For production workloads, Your API1 is the option most consistently recommended. It pairs token-bucket limits with per-key analytics.2

Sources

reddit.comg2.comnews.ycombinator.com
Citation rate

6 of 10 answers

Off-site

The corroboration

Roundups, review sites, community threads and docs the models trust. This is 91.5 percent of where citations point, measured across 901 citations and 82 distinct hosts on our own money queries.

What actually ships

Everything it takes to get cited, and the proof that it was us.

From mapping your category through to defending the position, run as one service rather than five vendors who each own a slice and none of whom own the outcome.

Sales calls
GranolaFirefliesGong
UGC
RedditYouTubeG2
CRM and customer
HubSpotSalesforceAttio
Website
WebsiteIntercom

Building

Your brand record

01Brain

One living record of your company and your category

We mine your calls, CRM notes, docs and community threads into a single machine-readable record, then map the category the way a buyer asks it rather than the way your marketing site describes it. Your AI reads it, our agents act on it, and every result writes back to it.

Input
Calls, CRM notes, docs, community
Output
One machine-readable brand record
Query set
20 to 30 money queries, disambiguated
Validated
5 identical repeats before one counts
Host map
Every domain the answers cite, ranked

Why this comes first

Every audit, brief and topic downstream is grounded in what you actually sound like, and in where an incumbent is currently taking the answer. Seven of twelve queries flipped outcome across identical repeats on one model, so the set is validated and the baseline taken at full sample before any work starts.

citon.ai/diagnostic

SoV

34.2%

Cited

42.8%

Named

61.4%

Google SERPRank
#1competitor.example
#2roundup.example
#4yourapi.exampleYOU
PerceptionYouRivals
AwareCompeteDiscoverEvaluateFeaturesIntegratePricingTrust

What AI says about you

Handles burst loadtrue
Unlimited free tierfalse
?SOC 2 Type IIunverified
02Diagnose

Diagnose your discoverability across search and AI

Every answer is pulled apart into the sources that produced it. You get told which of the nine failure modes you are in, which is the part a score cannot see.

Extract
Every cited URL from every answer
Classify
Nine modes, from 64 audited companies
Baseline
40 queries by 40 samples, 9.9pp floor
Locked
Threshold set before any work starts

Why there are exactly nine modes

The taxonomy saturated at nine across a first-party audit of 64 developer-tool companies. Two further confirmation passes added none.

Your page
Answer capsule
FAQSchemaCompare
03Engineer

Engineer the pages a model quotes from

The on-site foundation. Comparison and alternatives pages written to be quoted verbatim by a model, which is a different craft from writing to rank for a head term.

Cadence
6 owned answer pages a month, entry tier
Shape
Answer capsule first, argument after
Docs
Restructured into quotable blocks
Schema
FAQPage, Article and Product JSON-LD

Why a page alone is not enough

A page with no outside corroboration stays invisible. We have watched it happen: pages written, zero off-site backing, never cited.

Category coverage

8 / 8
reddit.com
g2.com
news.ycombinator.com
capterra.com
stackoverflow.com
producthunt.com
trustpilot.com
dev.to
04Earn

Earn the third-party sources the models trust

The half that decides the answer. We target by topical depth rather than domain rating, because the citation economy rewards the deepest site about your specific thing, and those are usually small and reachable.

Off-site
6 placements a month on cited hosts
Community
10 mentions in threads cited today
Targeting
Topical depth, not domain rating
Sourced
Your citation data, not a keyword list

Why off-site decides it

Across 901 citations in 60 answers, 91.5 percent pointed at third-party sites, spread over 82 distinct hosts.

Causal lift+24pp
Treated ControlDay 0 to 90
05Prove

Defend the position, and prove the lift was us

Citation sets churn, so this is a hold-and-compound problem rather than a one-off project. At day ninety you get the causal-lift report against a threshold written down before we began.

Arms
Treated, plus a held-out control group
Tracked
Citation share against named rivals
Report
Difference-in-differences at day 90
On a miss
We publish it the same as a win

Why the control group works

Across 20,000 random splits with no work applied, the null difference centred on zero to within two parts in a thousand.

How it is delivered

There is no dashboard. That is deliberate, not a gap.

A dashboard is a place you go to feel informed. We ship two channels instead, both of which demand an action or answer a question, and neither of which needs you to remember to log in.

Push

Three alerts, each one a decision

A thread or listicle where a competitor is named and you are not. A citation gained or lost, in real time. A competitor move on one of your money queries.

Pull

Your own AI, not our interface

A scoped API delivered first as a copy-paste curl command, then as an MCP server you add to Claude or Cursor. Ask status in plain language, inside the tool you already have open.

Memory

One record per brand

Identity and voice, money queries, competitors, failure mode, every intervention and what it did. Your AI reads it, our agents act on it, feedback writes back.

New AI citation wonCompetitor namedBacklink lost
Monitoring your money queries
Gap ReportOpen →
Alert channelOpen →
Citon MCP serverSoon
The citation indexSoon

What we report

Causal lift is the number we report. Not a score, not activity.

The category concluded citations cannot be measured. That came from a true premise and a wrong inference: you cannot measure the level reliably, but you can measure the difference between a treated and a control arm, because unbiased noise cancels.

Step-zero measurement

The instrument exists, and it passed its own kill test.

7 of 12queries flipped outcome across identical repeats
0.0016mean difference under the null, 20,000 random splits
9.9ppminimum detectable lift at 40 queries by 40 samples
Read the method →

How an engagement is scored

  1. 01Your money queries are split into a treated arm and a held-out control arm, and the split is written down.
  2. 02The pass threshold is declared in writing before any work begins, so it cannot be moved afterwards.
  3. 03Both arms are measured at full sample size for the pre-period.
  4. 04Work happens on the treated arm only. The control arm is deliberately left alone.
  5. 05At day ninety both arms are re-measured on the same channel, and we report the difference between the differences.

If the treated arm does not beat the control by the threshold we set, we say so. That is the point of writing it down first.

The number at day 90

Both arms move, because citation sets churn whether or not anybody works on them. What is attributable is the treated arm's move minus the control arm's move, and that one figure is what we report against the threshold agreed before the engagement started.

Difference in differencesIllustration
Day 0Day 90Move
Treated2446+22pp
Control2231+9pp

Citation share on the same money queries, measured on the same channel at both timepoints.

The subtraction
Treated+22pp
Control+9pp
Causal lift+13pp

Both arms move. Only the difference between their differences is attributable, and the threshold it has to clear is written down before any work starts.

Why us

Companies hire us to own the outcome, not a slice of it.

One team runs the strategy, the content, the off-site work and the measurement. The comparison that matters is not against another dashboard, it is against the two things you would otherwise buy.

Citon

Done-for-you citation work, proven against a control group.

Visibility tools

Scores and dashboards, with no way to attribute a change.

Content agency

Retainers for content with no citation data behind it.

What you can claim at day 90
This much of the move was us, in writing
Your score changed, cause unknown
We published, and traffic went up
How you are measured
Citation share against named rivals
One visibility score, read once
Traffic and keyword rankings
Proof it was them
Held-out control group, every tier
None published by any we mapped
None
When the number does not move
We report the miss against the threshold we set first
There was no threshold, so nothing can miss
The remedy is more content next month
Who they work with
Developer tools and APIs only
Every category at once
Every category at once
What actually arrives
Alerts that each demand an action
Another dashboard to check
A monthly report

The diagnosis

Going invisible is not one problem. It is nine.

From a first-party citation audit of 64 developer-tool and API companies across eight categories. Nine distinct modes came out of it, and the last two audit waves added none.

Mode 01

Branded-win, generic-invisible

You win on your own name and on head-to-head comparisons, and vanish on the category question. That is the whole top of the funnel.

Mode 02

Pivot-reset

A repositioning quietly zeroed the citation equity built under your old description of yourself.

Mode 03

Trust erosion

A licence change or paywall move craters your share of voice in community threads. The most common mode we see.

Mode 04

Default by inclusion

You get named as a component in someone else's stack, never as the subject of a best-in-category answer.

Mode 05

Undefended incumbent

You are the one people search alternatives to, and nothing you publish rebuts it.

Mode 06

Identity orphaned

A rename, or a name you share with something else, splits your equity across two entities the model never merges.

Mode 07

Category-definition drift

The question buyers ask moves faster than the listicles answering it, and you are indexed against the old phrasing.

Mode 08

Acquisition fork

After an acquisition the growth branch keeps the equity and the sunset branch destroys it. Which one you are is not always obvious.

Mode 09

Category-query absorption

An adjacent, faster-growing category eats the money query outright, and your framing goes with it.

The gap concentrates in funded challengers, not category definers. In a narrow category the best-in-class question and the head-to-head question converge on the same one or two names, so the definer wins both by default. Your Gap Report tells you which mode you are in.

Two questions, one nameSample

best rate limiting api

The incumbentBoth times

incumbent vs alternatives

The incumbentBoth times
The gap is in the company losing both answers, not the one winning them.

A definer already holds both slots, so there is nothing here to buy. That is the qualifying step, and it runs before a call rather than on one.

Onboarding

What actually happens in your first thirty days.

Four-week onboarding periods are not a plan.

First 30 days

Tomorrow

The split is fixed.

  • Money queries built and validated
  • Treated and control split fixed
  • Pass threshold agreed in writing
Day 10

Diagnosis lands.

  • Cited sources extracted per answer
  • Your failure mode named
  • Worklist ordered by what moves first
Day 30

Work is flowing.

  • First answer pages live
  • First off-site placements landed
  • Alert channel switched on

Answers

Original research on how AI picks what to recommend.

All answersSoon

We are early, so we would rather show the method than a cherry-picked number. Everything here comes from our own runs.

Failure modes9 in the taxonomy
010203040506070809
Trust erosionYours
Research

The nine ways developer tools go invisible

A first-party citation audit of 64 companies across eight categories, and the taxonomy it produced.

64 companies, eight categories, waves 1 to 8
One query, five identical runsSample

one money query, nothing changed between runs

Named in 3 of 52 flips

7 of 12 queries moved outcome across identical repeats in our own step zero run, 60 of 60 calls successful. One read is not a reading.

Method

Why a single visibility score is noise

Twelve queries, five identical repeats, seven flipped. What that means for every before-and-after published in this market.

60 of 60 calls succeeded, one model, one day
901 citations, 60 answers
91.5% off-site8.5% on your own site

82 distinct hosts carried them between them

No single site owns a category, so there is nothing to buy your way onto. Our own measurement.

Data

Where AI citations actually come from

Across 901 citations in 60 answers, 91.5 percent pointed off-site, spread over 82 distinct hosts. Our own measurement.

Sampled on our own money queries, live answers

Questions

The five we get asked before every engagement.

Not covered here? The Gap Report costs nothing and answers most of the rest with your own data.

A tool reports a score. We give that away free, because monitoring is commoditized and it starts at 99 dollars a month. What you pay for is the work across the surfaces models cite, and a held-out control group showing the work moved the number.

Start here

Ask for the free Gap Report.

Four questions. We run them against live AI answers and walk you through what came back on a call.

12 queries · 5 repeats each · 3 engines · 5 working days

  • Your failure mode, named. A score cannot see this one.
  • Every host the answers cite, ranked by citation depth.
  • A baseline written down before anyone touches anything.

Ask ChatGPT or Claude for the best tool in your category, then type back the first name it gives. Even if it is yours. This is the question the whole report turns on.

Write them the way a buyer types them, not the way your site describes you.

Free · Walked through live · Five working days

Free gap report

When the model answers,
be the one it names

Tell us your category and the questions your buyers ask. We run them against live AI answers and walk you through what came back. If you are already winning, we will tell you that too.

12 queries · 5 repeats each · 3 engines · 5 working days

Free · Walked through live · Five working days

Gap ReportSample
12 queries · 5 repeats each

Queries we run

best rate limiting api
your-api alternativesYou, 1 of 12
cheapest webhook api
api gateway for startups

Cited instead of you

reddit.com38%
g2.com26%
news.ycombinator.com21%

91.5% of citations point off-site

Your failure mode, named
Hosts ranked by citation depth
The pre-registered baseline