For AI infrastructure, developer tools, APIs, CLIs and MCP servers
Your buyers asked an AI which tool to use. It named your competitor.
We measure where you stand in live AI answers, we show you the exact pages the models cite when your buyers ask, and we go get you onto them. Every engagement runs against a Why this is the only number that settles the argument.The baseline is registered before any work starts, and the treated and held-back queries are split at the same moment. At day 90 the reading is a comparison against a group we deliberately did not touch. Without that split, a lift claim is either unfalsifiable or an unbounded liability., so at day 90 you get the one number nobody else in this category will give you: how much of the movement we caused.
Your buyer asks
The answer they get
For production workloads, most teams land on Competitor API1. It pairs token-bucket limits with per-key analytics.2
Where the citations resolve
Measured across
91.5%
of AI citations point off-site
901 citations across 60 answers. Our own measurement.
82 hosts
carried those citations between them
No single site owns a category, so there is nothing to buy your way onto.
45.5%
of citations change between consecutive observations
Ahrefs, 43,000 keywords. AI Overviews persist 2.15 days.
The AI-citation team for developer tools
We run one play, for one kind of company, all the way down.
Citon works with companies building AI infrastructure, developer tools, APIs, CLIs and MCP servers. Not general B2B SaaS. The surfaces a developer evaluates from are a short and specific list, and it is the same list every time, which is why we work one niche rather than several.
- One niche
- Held-out control, every tier
- A few clients per quarter
- No dashboard at any tier
A small senior team. No junior handoff after the pitch.
Where developers decide
GitHub and Hacker News
Where a tool gets vouched for
Docs as answers
Usually the surface that does the most work per hour spent on it
Best-in-category roundups
Ranked by topical depth, never by domain rating
Community threads
Already being cited, so already worth entering
Review sites
The corroboration models trust
Not a marketing blog post
Short answer
How is this different from an AI visibility tool?
A tool reports a score. We give that away free, because monitoring is commoditized and it starts at 99 dollars a month. What you pay for is the work across the surfaces models cite, and a held-out control group showing the work moved the number.
The problem with the current answer
You already know AI decides. Every vendor still sells you a number that moves on its own.
You see a score. Nobody tells you why.
Level measurement is destroyed by non-determinism. We ran twelve queries five times each on one model and seven flipped outcome. That is a property of the system, not a vendor failure.
Your agency is a step behind on AI.
You get a ranking report every month while the clicks behind those rankings fall away. The surfaces that decide an AI answer are mostly not the ones an SEO retainer works on.
Reports tell you what happened. Not what you caused.
Across more than fifty companies mapped in this category, we have not found one publishing a control group or a causal-lift number.
Monitoring is commoditized. It starts at 99 dollars a month, and Google ships AI-search data free inside Search Console. So we give monitoring away and charge for the part nobody else will stand behind.
What changed
Developers stopped browsing results. They ask, and they take the answer.
We run on-site and off-site as one service, because a page with no outside corroboration stays invisible, and corroboration with no page has nothing to point at.
Rate limiting for production APIs
competitor.example › blog
roundup.example › guides
On-site
The citable answer
The page the model points at. Necessary, and table stakes. We do this competently and we do not try to out-automate the page engines that already do it at scale.
Which API should I use for rate limiting?
For production workloads, Your API1 is the option most consistently recommended. It pairs token-bucket limits with per-key analytics.2
Sources
6 of 10 answers
Off-site
The corroboration
Roundups, review sites, community threads and docs the models trust. This is 91.5 percent of where citations point, measured across 901 citations and 82 distinct hosts on our own money queries.
What actually ships
Everything it takes to get cited, and the proof that it was us.
From mapping your category through to defending the position, run as one service rather than five vendors who each own a slice and none of whom own the outcome.
Building
Your brand record
One living record of your company and your category
We mine your calls, CRM notes, docs and community threads into a single machine-readable record, then map the category the way a buyer asks it rather than the way your marketing site describes it. Your AI reads it, our agents act on it, and every result writes back to it.
- Input
- Calls, CRM notes, docs, community
- Output
- One machine-readable brand record
- Query set
- 20 to 30 money queries, disambiguated
- Validated
- 5 identical repeats before one counts
- Host map
- Every domain the answers cite, ranked
Why this comes first
Every audit, brief and topic downstream is grounded in what you actually sound like, and in where an incumbent is currently taking the answer. Seven of twelve queries flipped outcome across identical repeats on one model, so the set is validated and the baseline taken at full sample before any work starts.
SoV
34.2%
Cited
42.8%
Named
61.4%
What AI says about you
Diagnose your discoverability across search and AI
Every answer is pulled apart into the sources that produced it. You get told which of the nine failure modes you are in, which is the part a score cannot see.
- Extract
- Every cited URL from every answer
- Classify
- Nine modes, from 64 audited companies
- Baseline
- 40 queries by 40 samples, 9.9pp floor
- Locked
- Threshold set before any work starts
Why there are exactly nine modes
The taxonomy saturated at nine across a first-party audit of 64 developer-tool companies. Two further confirmation passes added none.
Engineer the pages a model quotes from
The on-site foundation. Comparison and alternatives pages written to be quoted verbatim by a model, which is a different craft from writing to rank for a head term.
- Cadence
- 6 owned answer pages a month, entry tier
- Shape
- Answer capsule first, argument after
- Docs
- Restructured into quotable blocks
- Schema
- FAQPage, Article and Product JSON-LD
Why a page alone is not enough
A page with no outside corroboration stays invisible. We have watched it happen: pages written, zero off-site backing, never cited.
Category coverage
Earn the third-party sources the models trust
The half that decides the answer. We target by topical depth rather than domain rating, because the citation economy rewards the deepest site about your specific thing, and those are usually small and reachable.
- Off-site
- 6 placements a month on cited hosts
- Community
- 10 mentions in threads cited today
- Targeting
- Topical depth, not domain rating
- Sourced
- Your citation data, not a keyword list
Why off-site decides it
Across 901 citations in 60 answers, 91.5 percent pointed at third-party sites, spread over 82 distinct hosts.
Defend the position, and prove the lift was us
Citation sets churn, so this is a hold-and-compound problem rather than a one-off project. At day ninety you get the causal-lift report against a threshold written down before we began.
- Arms
- Treated, plus a held-out control group
- Tracked
- Citation share against named rivals
- Report
- Difference-in-differences at day 90
- On a miss
- We publish it the same as a win
Why the control group works
Across 20,000 random splits with no work applied, the null difference centred on zero to within two parts in a thousand.
How it is delivered
There is no dashboard. That is deliberate, not a gap.
A dashboard is a place you go to feel informed. We ship two channels instead, both of which demand an action or answer a question, and neither of which needs you to remember to log in.
Three alerts, each one a decision
A thread or listicle where a competitor is named and you are not. A citation gained or lost, in real time. A competitor move on one of your money queries.
Your own AI, not our interface
A scoped API delivered first as a copy-paste curl command, then as an MCP server you add to Claude or Cursor. Ask status in plain language, inside the tool you already have open.
One record per brand
Identity and voice, money queries, competitors, failure mode, every intervention and what it did. Your AI reads it, our agents act on it, feedback writes back.
What we report
Causal lift is the number we report. Not a score, not activity.
The category concluded citations cannot be measured. That came from a true premise and a wrong inference: you cannot measure the level reliably, but you can measure the difference between a treated and a control arm, because unbiased noise cancels.
Step-zero measurement
The instrument exists, and it passed its own kill test.
How an engagement is scored
- 01Your money queries are split into a treated arm and a held-out control arm, and the split is written down.
- 02The pass threshold is declared in writing before any work begins, so it cannot be moved afterwards.
- 03Both arms are measured at full sample size for the pre-period.
- 04Work happens on the treated arm only. The control arm is deliberately left alone.
- 05At day ninety both arms are re-measured on the same channel, and we report the difference between the differences.
If the treated arm does not beat the control by the threshold we set, we say so. That is the point of writing it down first.
The number at day 90
Both arms move, because citation sets churn whether or not anybody works on them. What is attributable is the treated arm's move minus the control arm's move, and that one figure is what we report against the threshold agreed before the engagement started.
Citation share on the same money queries, measured on the same channel at both timepoints.
Both arms move. Only the difference between their differences is attributable, and the threshold it has to clear is written down before any work starts.
Why us
Companies hire us to own the outcome, not a slice of it.
One team runs the strategy, the content, the off-site work and the measurement. The comparison that matters is not against another dashboard, it is against the two things you would otherwise buy.
Citon
Done-for-you citation work, proven against a control group.
Visibility tools
Scores and dashboards, with no way to attribute a change.
Content agency
Retainers for content with no citation data behind it.
The diagnosis
Going invisible is not one problem. It is nine.
From a first-party citation audit of 64 developer-tool and API companies across eight categories. Nine distinct modes came out of it, and the last two audit waves added none.
Branded-win, generic-invisible
You win on your own name and on head-to-head comparisons, and vanish on the category question. That is the whole top of the funnel.
Pivot-reset
A repositioning quietly zeroed the citation equity built under your old description of yourself.
Trust erosion
A licence change or paywall move craters your share of voice in community threads. The most common mode we see.
Default by inclusion
You get named as a component in someone else's stack, never as the subject of a best-in-category answer.
Undefended incumbent
You are the one people search alternatives to, and nothing you publish rebuts it.
Identity orphaned
A rename, or a name you share with something else, splits your equity across two entities the model never merges.
Category-definition drift
The question buyers ask moves faster than the listicles answering it, and you are indexed against the old phrasing.
Acquisition fork
After an acquisition the growth branch keeps the equity and the sunset branch destroys it. Which one you are is not always obvious.
Category-query absorption
An adjacent, faster-growing category eats the money query outright, and your framing goes with it.
The gap concentrates in funded challengers, not category definers. In a narrow category the best-in-class question and the head-to-head question converge on the same one or two names, so the definer wins both by default. Your Gap Report tells you which mode you are in.
best rate limiting api
incumbent vs alternatives
A definer already holds both slots, so there is nothing here to buy. That is the qualifying step, and it runs before a call rather than on one.
Onboarding
What actually happens in your first thirty days.
Four-week onboarding periods are not a plan.
First 30 days
The split is fixed.
- Money queries built and validated
- Treated and control split fixed
- Pass threshold agreed in writing
Diagnosis lands.
- Cited sources extracted per answer
- Your failure mode named
- Worklist ordered by what moves first
Work is flowing.
- First answer pages live
- First off-site placements landed
- Alert channel switched on
Answers
Original research on how AI picks what to recommend.
We are early, so we would rather show the method than a cherry-picked number. Everything here comes from our own runs.
The nine ways developer tools go invisible
A first-party citation audit of 64 companies across eight categories, and the taxonomy it produced.
64 companies, eight categories, waves 1 to 8one money query, nothing changed between runs
7 of 12 queries moved outcome across identical repeats in our own step zero run, 60 of 60 calls successful. One read is not a reading.
Why a single visibility score is noise
Twelve queries, five identical repeats, seven flipped. What that means for every before-and-after published in this market.
60 of 60 calls succeeded, one model, one day82 distinct hosts carried them between them
No single site owns a category, so there is nothing to buy your way onto. Our own measurement.
Where AI citations actually come from
Across 901 citations in 60 answers, 91.5 percent pointed off-site, spread over 82 distinct hosts. Our own measurement.
Sampled on our own money queries, live answersQuestions
The five we get asked before every engagement.
Not covered here? The Gap Report costs nothing and answers most of the rest with your own data.
A tool reports a score. We give that away free, because monitoring is commoditized and it starts at 99 dollars a month. What you pay for is the work across the surfaces models cite, and a held-out control group showing the work moved the number.
Start here
Ask for the free Gap Report.
Four questions. We run them against live AI answers and walk you through what came back on a call.
12 queries · 5 repeats each · 3 engines · 5 working days
- Your failure mode, named. A score cannot see this one.
- Every host the answers cite, ranked by citation depth.
- A baseline written down before anyone touches anything.
Free gap report
When the model answers,
be the one it names
Tell us your category and the questions your buyers ask. We run them against live AI answers and walk you through what came back. If you are already winning, we will tell you that too.
12 queries · 5 repeats each · 3 engines · 5 working days
Free · Walked through live · Five working days
Queries we run
Cited instead of you
91.5% of citations point off-site