This guide, updated October 2026, explains why the apps and the APIs answer differently and how to track the answers buyers read. Far & Wide asks buyer questions in the real consumer web interfaces of ChatGPT, Gemini, Microsoft Copilot, Perplexity and Google AI Mode, with web search on, logged out, from the brand’s own country, and reads Google AI Overviews from the live Google results page.
TL;DR
- Check in the apps, not the API. In OpenAI’s API, web search and location are off unless the developer sets them (OpenAI API docs); the ChatGPT app searches, shows sources and localises by IP on its own (OpenAI Help Center).
- Check logged out. Your own account can carry memory and past chats.
- Ask more than once. In a SparkToro study, there was less than a 1 in 100 chance that ChatGPT or Google’s AI would give the same list of brands twice if asked 100 times (SparkToro).
- Read your position as a rate. “Named in 7 of 10 runs” tells you more than any single answer.
How can you tell what ChatGPT says about your company?
You can tell what ChatGPT says about your company by asking the questions your buyers ask in the ChatGPT app the way a stranger would: logged out, with web search on, from your buyers’ country, and several times per question. Then note whether you were named, where, what it claimed and which sources it showed.
Your own logged-in account is a poor test. When Memory is on, ChatGPT “can remember relevant preferences and details from your chats” (OpenAI Help Center). OpenAI began letting people use ChatGPT without an account in 2024, starting in a few markets and rolling out gradually (TechCrunch), and that logged-out view is closer to what a new buyer gets.
Location also changes what ChatGPT shows. ChatGPT “may use your IP address to estimate your general location, such as your country, state, or city”, and a VPN can change that estimate (OpenAI Help Center).
A manual check, step by step
- Write down 5 to 10 questions your buyers ask, in their words (how to find them).
- Log out of ChatGPT, or open a private browser window where you are not signed in.
- Check from your buyers’ country. If you sell abroad, use a connection in that country.
- Let ChatGPT search the web. When it does, a Sources button shows the pages it used (OpenAI Help Center).
- Ask each question at least three times in fresh chats, because the brand list changes between runs.
- Log every answer: named or not, position, claims, competitors named, sources shown.
- Repeat in the other apps your buyers use, such as Gemini, Microsoft Copilot, Perplexity and Google AI Mode.
For a fuller walkthrough of the hand check, see how to check what ChatGPT says about your brand.
Reading what you find
Each answer usually falls into one of four groups.
| What you see | What it usually means | What to check next |
|---|---|---|
| Absent | The answer draws on pages that don’t mention you | Open Sources and read the pages that won |
| Present but wrong | A page the app read says something wrong about you | Find the cited page, then fix the wrong information |
| Present but behind competitors | Others appear on more of the pages the app reads | Compare which sources name them and not you |
| Present and cited | Your own page is feeding the answer | Keep that page accurate and current |
A mention means the answer names your brand. A citation means the answer links one of your pages as a source. You can get either one without the other, so record both.
A wrong claim that repeats in 2 of your 3 runs is worth tracing to its source; one that shows up only once may be a one-off.
Common mistakes when checking by hand
- Checking while logged in. Memory and past chats can shape the answer you get.
- Asking once. The brand list changes between runs, so one answer can mislead you.
- Using a developer API as a stand-in. Web search and location are off unless the developer turns them on.
- Checking from the wrong country. The app estimates your location from your IP address, so a VPN or a trip abroad changes it.
Why do the ChatGPT app and the API give different answers?
The ChatGPT app and the API give different answers because the API is a building block for developers and the app is a finished product with its own search, sources, location handling, memory and model.
| Real consumer app (what buyers use) | Developer API (what many scripts and tools call) | |
|---|---|---|
| Web search | ChatGPT “may search the web automatically” (OpenAI Help Center) | Off unless the developer adds the search tool or a search model (OpenAI API docs); Gemini’s API needs the google_search tool (Google AI for Developers) |
| Sources | A Sources view with cited links | Only when search is on; display is up to the developer |
| Location | Estimated from your IP address | Used only if the developer sends a country, city, region or timezone (OpenAI API docs) |
| Personalisation | Logged-in ChatGPT can use memory; signed-in AI Mode can use Search history | None, unless the developer sends context |
| Model | The model ChatGPT runs right now | Whatever model the developer picks |
Even the model can differ. OpenAI keeps a separate API alias, chat-latest, that “points to the latest Instant model currently used in ChatGPT”, and on the same page recommends a different model for production API use (OpenAI API docs). A script that follows that advice is not even asking the model ChatGPT runs.
Perplexity and Google split between app and API in the same way. Perplexity offers a developer API with “web-grounded answers with built-in citations” (Perplexity docs), but it is a separate product configured by whoever calls it. Google AI Overviews and AI Mode live inside Google Search, so to see them you look at Search itself. For how the assistants differ from each other, see our Perplexity vs ChatGPT vs Gemini comparison.
Why does the same question get a different answer each time?
The same question gets a different answer each time because these assistants generate a fresh answer on every run, and the brand list moves with it. In SparkToro’s study, 600 volunteers ran 12 prompts through three AI tools a combined 2,961 times, each keeping their usual settings, personalised or default (SparkToro). The chance that ChatGPT or Google’s AI would return the same list in 100 runs was under 1 in 100, and two lists in the same order were closer to 1 in 1,000.
Read visibility as a rate. SparkToro concluded that visibility percentage across many prompts, run several times, “is a reasonable metric”. In its data, one Los Angeles hospital showed up in 69 of 71 answers, a 97% visibility rate (SparkToro).
A second study, from Detailed, checked more than 1,300 prompts daily for four weeks, over 70,000 responses, in Google AI Overviews and ChatGPT. It did not cover Gemini, Copilot or Perplexity, and each prompt was checked once a day from the US without logging in. Its findings:
- The leading brand appeared on 89% of days in AI Overviews and 82% in ChatGPT.
- Comparing any two days, the brand set was identical on 1.1% of pairs in AI Overviews and 0.3% in ChatGPT.
- Day to day, core brands changed by 11% in AI Overviews and 13% in ChatGPT; tail brands by 65% and 78%.
In both studies, a category leader usually appears, while a brand at the edge of the list can look in or out by chance in any single check.
Is checking AI mentions by hand every week enough, or do you need automated tracking?
Checking AI mentions by hand every week is enough for a handful of questions on one or two apps, but it stops being repeatable once you need several runs per question across five or six AI surfaces. Automated tracking only beats it if the tool asks where buyers do: the consumer apps of ChatGPT, Gemini, Microsoft Copilot and Perplexity, plus Google AI Overviews and AI Mode on the live results page, logged out, with web search on, from your buyers’ country.
The arithmetic is easy to check. Ten buyer questions, five surfaces and three runs each make 10 × 5 × 3 = 150 answers per check, or about 600 a month if you check weekly. Time one answer yourself; at two minutes each, 150 answers take five hours a week.
The bigger problem with weekly hand checks is consistency, not time. Hand checks drift: one week you forget to log out, the next you check from abroad, and the change you see comes from your setup.
| Manual weekly check | Automated tracking in the real apps | |
|---|---|---|
| Answers per check | Whatever you have time for | The same number every time |
| Runs per question | Often one | Several, on a fixed schedule |
| Surfaces | Usually ChatGPT, sometimes one more | All the surfaces you choose |
| Conditions (login, search, country) | Depend on where you are and what you remember | Fixed in the setup |
| Best for | A first look, testing one change | Trends, share of voice, catching drops early |
How do you track whether AI is recommending you over time?
You track whether AI is recommending you over time by asking a fixed set of buyer questions on the same AI surfaces, under the same conditions, several times each, on a fixed schedule, and comparing the rate at which you are named from one period to the next.
AI visibility is how often, where and how accurately your brand appears in the answers AI assistants give your buyers, measured in the apps they use and read as a rate across many runs.
- Fix the question set. Use the questions buyers ask, and keep them the same between periods.
- Pick the surfaces your buyers use. In a Pew Research Center survey of U.S. adults, 44% said they use ChatGPT, 24% Gemini and 17% Copilot (Pew Research Center).
- Ask in the apps: logged out, web search on, from your buyers’ country.
- Run each question several times, because single answers swing.
- Record every answer with the fields in the checklist below.
- Repeat on a fixed schedule, such as every 3 or 7 days.
- Compare rates, not single answers: the share of runs that name you, your average position and your share of voice against competitors, and flag any drop in naming rate from one period to the next.
AI Overviews and AI Mode need their own check
Google says AI Overviews and AI Mode “may use different models and techniques, so the set of responses and links they show will vary”, and both may use “query fan-out”, running several related searches to build one answer (Google Search Central). AI Mode can also be personalised for signed-in users who have Search history and personalised recommendations turned on (Google Search Help). The brands shown also move from day to day: in Detailed’s four-week study, the set of brands in AI Overviews matched between any two days only 1.1% of the time (Detailed). So check both on the live results page, signed out, in your buyers’ country. Our Google AI Mode optimization guide covers what to change.
What every tracking record should store
Spreadsheet or tool, each answer should carry these 13 fields:
- Date and time of the run
- Platform (ChatGPT, Gemini, Copilot, Perplexity, AI Overviews, AI Mode)
- Interface: consumer app or developer API
- Logged in or logged out
- Web search on or off
- Country and language
- Exact question text and run number
- Brand named: yes or no, and position
- Competitors named
- Claims made about you
- Tone of the answer about you: positive, neutral or negative
- Sources shown, with URLs
- The full answer text
The interface, login, search and country fields explain most mismatches between two results.
Which tool is best for tracking how ChatGPT talks about your brand?
The best tool to track how ChatGPT talks about your brand is one that reads the same answers your buyers read: in the real ChatGPT app, with web search on, logged out, from your buyers’ country, several times per question on a schedule. It should also tell you whenever an answer came from a developer API instead of the app.
Tracking how ChatGPT talks about you comes down to four things per answer: whether you are named, where, what it says about you, and which sources it shows. The same applies to the other surfaces buyers use. About half of U.S. adults (49%) now use AI chatbots, according to Pew Research Center, and those are people using the chatbots themselves, not developers calling an API.
Share of voice shows how you compare with competitors on the same question set. Count every brand mention across all runs: if brands are named 200 times in total and you are named 40 of those times, your share of voice is 20%. The figure is only comparable over time if it is collected the same way every period.
Seven questions to ask any tracking tool
- Does it ask in the consumer app or through a developer API?
- Is web search on, and does it save the sources shown?
- Does it ask logged in or logged out?
- From which country, and in which language?
- How many runs per question, and how often?
- Are any API answers labelled as API answers?
- Which surfaces does it cover (ChatGPT, Gemini, Microsoft Copilot, Perplexity, Google AI Overviews, Google AI Mode), and do they match where your buyers are?
A tool that can’t answer the first question may be measuring something other than what your buyers see. For the metrics themselves, see what AI visibility tracking is.
FAQ
The tool and you may be asking under different conditions: app versus API, logged in versus logged out, search on versus off, or a different country. Answers also change from run to run, so compare rates over several runs, not one screen against one chart.
See the answers your buyers see
The free scan asks each of your buyer questions once, so you can see which questions you win and which you miss. Paid plans ask each question several times on a schedule, every 3 days 3 times or every 7 days 5 times; see pricing.
Start your free scan