What the AI Presence tab shows today, and how we make it richer using the full DataForSEO AI Optimization API — real answers from ChatGPT, Claude, Gemini and Perplexity, real history, and real AI search demand.
Read-only tab on the site detail page. Five sections, every value backed by real data (no placeholders). Full-access gated to Pro/Business/pilot; everyone else sees an upgrade CTA.
llm_mentions/historical/live), page-level citations (llm_mentions/top_mentioned_pages/live), and the monthly Live AI Presence sweep, which actually asks ChatGPT, Claude, Gemini and Perplexity the site's questions and records mention and citation per engine. Read the gap list below as the state before that work, not as what is missing now.| Section | What the customer sees | Data behind it |
|---|---|---|
| AI mentions over time | Hero count ("N this month"), "+P% vs last month", Google AI / ChatGPT split, a trend line | mentionsHistory (self-built, one point per refresh) + aiMentionMetrics |
| What Scaup did this week | Timeline of completed plan items | Plan items + stored executionSummary.what |
| Real AI answer examples | Question + the actual AI answer snippet | aiMentionExamples[].answerSnippet |
| You vs competitors, question by question | Questions where you're cited vs where only competitors show up + the plan to close each gap | aiVisibility.opportunities + competitorMentions |
| Coming up next | The single highest-volume gap Scaup will close next | Top opportunity by AI search volume |
Call (worker aiVisibility.ts) | Gives us |
|---|---|
llm_mentions/cross_aggregated_metrics/live | Mentions for you + up to 5 competitors; your Google/ChatGPT split (the hero number) |
llm_mentions/top_domains/live | Which domains AI pulls from in your niche |
llm_mentions/search/live | Individual Q&A records → opportunities where a competitor is cited but you're not |
Fetched by the worker during plan generation, competitor refresh, and a standalone ai_visibility_refresh job; stored on sites.aiVisibility / aiMentionMetrics / mentionsHistory.
The mentions data covers Google AI Overviews + ChatGPT only. Gemini and Perplexity never appear — they aren't mentions platforms and can only be reached via live queries.
History is built one point per refresh, so a new site has no chart for weeks. We could show a real multi-month history from day one.
We never show how many people actually ask AI about the customer's topic — only how often they're mentioned.
We infer mentions from a database. We never actually ask the models a live question and show "here's what ChatGPT says about you right now."
We show which domains get cited, not which specific pages — so we can't say "your pricing page is what AI quotes."
No view of which brand names AI associates with the niche.
Four product families under DataForSEO AI Optimization. Full endpoint/field reference lives in docs/dataforseo-ai-optimization-api.md.
| Product | What it does | Cost / call* | Status vs us |
|---|---|---|---|
| LLM Mentions | Brand/domain/keyword mentions inside AI answers, aggregated + individual + historical | ~$0.10 | using 3 of ~15 |
| AI Keyword Data | AI-platform search volume + monthly trend, per keyword (up to 1000/call) | low | unused |
| LLM Responses | Ask Claude / ChatGPT / Gemini / Perplexity a real prompt, get the real answer + citations | $0.008–0.55 | unused |
| LLM Scraper | Scrape ChatGPT / Gemini consumer UI for a keyword — SERP-style sources + brand entities | ~$0.004 | unused |
*Rough per-call figures from the docs' sample responses. LLM Responses is token-priced — cheap models (Haiku, gpt-4o-mini, sonar) are ~1¢, Opus with web search can hit 55¢.
historical/live (real monthly time series back to 2025-08),
target_metrics/live, top_mentioned_pages/live,
top_mentioned_brands/live, timeseries_delta & timeseries_new_lost.
The API exposes engines through two different data paths, and they don't cover the same engines. This is the key constraint behind "show all 4 engines."
| Engine | Mentions & history aggregated "who AI recommends," from a database |
Live LLM Responses we ask it a real prompt, see if you're cited |
|---|---|---|
| ChatGPT | ✓ platform: chat_gpt | ✓ |
| Google AI Overviews | ✓ platform: google | ✓ (via Gemini) |
| Gemini | ✗ not a mentions platform | ✓ |
| Perplexity | ✗ not a mentions platform | ✓ (Live only, no async) |
Three phases, ordered by impact-per-effort. Phase 1 upgrades the tab we already have; Phases 2–3 add new proof and new surfaces.
historical/live. Replace the sparse self-built mentionsHistory with a true monthly series (mentions + AI search volume) back to Aug 2025. A brand-new customer sees a full chart on day one instead of waiting weeks. One call, ~$0.10, per refresh.search / top_domains / cross_aggregated_metrics; the current published API uses search_mentions / top_mentioned_domains / target_metrics. Confirm ours are still-supported aliases or migrate — cheap insurance against a silent all-zero tab.postback_url, and cache — don't re-ask every page load. Perplexity is Live-only.top_mentioned_pages/live). Move from domain-level to page-level citations — lets us say "your guide page is what AI pulls from" and target content accordingly.top_mentioned_brands / brand_categories). Which brands AI ties to the niche — competitive-landscape view.timeseries_delta / timeseries_new_lost). "You gained N new AI mentions this month" — positive, movement-only framing that fits the product rule.Our customers care about GEO (getting recommended by AI engines) as much as classic SEO. These build on the same API — most reuse fields we already pay for — and turn the tab from "measuring" into "measuring and doing." Slot them alongside Phase 2.
fan_out_queries — the sub-questions the model expands a query into. These are the literal prompts to write content for. Turn them into new_content plan items. Pure GEO signal, zero extra API cost.top_mentioned_domains / top_mentioned_pages shows which sources AI pulls from in the niche (Reddit, Wikipedia, directories, specific competitors). Reframe from a stat into a tactic: "AI trusts these places — here's where to get you listed or mentioned."brand_entities + the live answers reveal what AI actually says about the customer — attributes, category, facts. Flag when AI is wrong (wrong pricing, "permanently closed", stale info). Correcting AI hallucinations about your own brand is one of the biggest GEO wins and nobody else surfaces it.target_metrics/live gives your mentions vs the niche total → "you appear in ~12% of AI answers here." Never a rank — a share that grows.llms.txt — the technical factors that make a page easy for an AI to parse and cite. Surface as a plain-language checklist Scaup works through; we could even auto-generate and maintain llms.txt.Exactly how the redesigned AI tab lays out, top to bottom. Left = a rough mock of what the customer sees; right = what it shows and where the data comes from. Same rules as today: full access gated to Pro/Business/pilot, and each section renders only when real data exists.
Recommended by 3 of 4 AI assistants this week.
The headline. One badge per assistant — cited (named + linked in the answer), mentioned (named, not linked), or not yet. A plain summary line: "recommended by 3 of 4." Never a rank; the number only ever grows as we win engines.
Data: live LLM Responses — we ask each of ChatGPT, Claude, Gemini, Perplexity the customer's key questions and check if the site is cited. The only view that covers all four. GEO
Google AI 30 · ChatGPT 12
The count of how often AI answers mention the site, as a hero number with month-over-month growth and a real trend line — full history from day one (no more waiting weeks for the chart to fill). Split by Google AI vs ChatGPT.
Data: LLM Mentions historical/live. Scope note: Google + ChatGPT only — the mentions database has no Gemini/Perplexity data, so this section is explicitly labeled as the 2-engine deep view (Section 1 is the all-4 view). SEO + GEO
Real demand: how many people ask AI about the customer's topics each month, and which questions are rising. Reframes the tab from "who mentions you" to "here's the demand, and here's your slice of it."
Data: AI Keyword Data keywords_search_volume/live (AI search volume + monthly trend, batched up to 1000 keywords). SEO + GEO
"best tattoo studio in Herzliya"
"…among the top-rated studios, Costello Tattoo stands out for fine-line work…"
The proof. For each key question, the real answer each assistant gave, with the mention highlighted and whether the site was cited. Expandable per engine. This is the strongest demo moment — "here's literally what ChatGPT tells your customers right now."
Data: live LLM Responses (all 4 engines). Cost-guarded: capped questions, cheap models, async task flow, cached. GEO
The merged list of questions where you're cited vs where only competitors show up, each with the plan to close the gap. Extended to all engines via the live checks, not just Google + ChatGPT.
Data: LLM Mentions opportunities (existing) + live LLM Responses per competitor. GEO
We'll work on getting you seen on these.
The sources AI trusts in this niche — the sites it pulls answers from. Reframed as a tactic, not a stat: "AI leans on these places, here's where to get you listed or mentioned." Drives concrete GEO plan items.
Data: LLM Mentions top_mentioned_domains / top_mentioned_pages. GEO
Flags when an assistant states something wrong about the business — wrong hours, wrong price, "permanently closed," outdated info. Correcting AI hallucinations about your own brand is a top GEO win and something no competitor tool surfaces. Only shows when there's something to fix.
Data: brand_entities + the live answer text, checked against known site facts. GEO
The activity timeline, now closing the loop: after a change ships, a later live re-check shows the result — "after we added this FAQ, ChatGPT now cites you for X." The before/after proof that ties our work to real movement in AI answers. Ideal for the weekly email.
Data: completed plan items + executionSummary + a scheduled live re-query after apply. SEO + GEO
The on-site GEO checklist — the technical things that make pages easy for AI to read and cite (schema, FAQ markup, clean content, llms.txt). Shown as plain-language items Scaup works through, not a technical audit. The "doing," paired with all the "measuring" above.
Data: our own site scan + JSON-LD work (existing) + auto-generated llms.txt. GEO
| New dataset | Feeds | Phase |
|---|---|---|
LLM Mentions historical/live | "AI mentions over time" trend line — Google + ChatGPT only (labeled as such) | 1 |
| LLM Responses (ChatGPT, Claude, Gemini, Perplexity) | Primary all-4-engines display — one row per engine, "what every AI says about you" | 2 |
| AI Keyword Data | New "AI demand for your topic" section | 2 |
top_mentioned_pages | Upgrades citations from domain- to page-level | 2 |
| LLM Scraper | Real answer snapshots / screenshots | 3 |
| Brands / timeseries | Competitive intel + momentum widgets | 3 |
fan_out_queries (already returned) | GEO content targets → new_content plan items | GEO |
brand_entities + live answers | Brand-fact monitoring — flag wrong facts AI states about you | GEO |
| Scheduled live queries | Question coverage tracker + before/after proof loop | GEO |
On-site: schema / FAQ / llms.txt | GEO-readiness checklist Scaup works through | GEO |
Cadence is set by cost, not by how often the data could change. The cheap database calls run weekly so trends stay alive; the token-priced live-model calls run monthly (or on demand) so cost stays flat per site. Every section carries a plain "last checked" line so the customer always knows how current the number is.
| Data source | Feeds section | Cadence | Why |
|---|---|---|---|
LLM Mentions historical/live | 2 · mentions trend | Weekly | ~$0.10/call. Cheap enough to run with the weekly report; keeps the trend line moving. |
LLM Mentions target_metrics / top_mentioned_* | 5 · competitors, 6 · sources | Weekly | Same low-cost database path; refresh alongside the mentions trend. |
AI Keyword Data keywords_search_volume | 3 · AI demand | Monthly | Volume is a monthly figure at source — re-pulling weekly shows the same number. Aligns with the monthly plan refresh. |
| LLM Responses (all 4 engines) | 1 · hero, 4 · live answers | Monthly + on-demand | Token-priced and up to 120s each — the one expensive call. Monthly sweep per site, plus a live re-check the customer can trigger. |
| LLM Responses (before/after re-query) | 8 · proof loop | Event-driven | Fires once, a few days after a fix ships — not on a clock. Cost scales with work done, not with time. |
Brand-fact check (brand_entities + live text) | 7 · accuracy | Monthly | Rides on the monthly live sweep — no extra call. Wrong-fact flags surface as soon as the sweep sees them. |
| LLM Scraper / brands / timeseries | Phase 3 depth | Monthly | Nice-to-have richness; monthly keeps cost negligible. |
On-site GEO scan (schema / FAQ / llms.txt) | 9 · readiness | Weekly | Our own scan, no API cost — runs with the existing weekly site check. |
Because sections refresh on different clocks, the customer must never guess how current a number is. Each section shows its own "last checked" line, and the tab caches results rather than re-querying on load.
lastCheckedAt per data source on the site record.historical/live swap. It's a one-call backend change, instantly fixes the weakest part of the tab (the empty trend line), and needs no new UI. Then prototype the Phase 2 LLM Responses section on a single site to size real cost before rolling out.