Why do it yourself first
Before spending on tools or agencies, there is real value in testing your AI visibility by hand. It is free, it takes an hour, and it gives you something no dashboard can: the actual words AI uses to describe you, in the answers your customers see.
Doing it yourself also builds intuition. When you run the prompts personally and watch what appears, you develop a feel for how AI answers your category, who it favours, and where you stand. That understanding makes every later decision sharper, whether you eventually buy a tool or not, and it costs nothing but attention.
There is a confidence benefit too. It is easy to feel anxious and powerless about AI visibility when it is an abstract worry, but running the test turns that feeling into information you can act on.
You may discover you are in better shape than you feared, or you may confirm a real gap, but either way you are no longer guessing. That shift from vague dread to concrete findings is worth the hour on its own, and it is why we start every engagement with exactly this kind of manual look before recommending anything.
There is a reason to start now rather than later. With AI adoption rising fast, as Pew Research documents in its work on how people are using AI in daily life, the answers being formed about you today are already shaping decisions. A manual audit is the fastest way to see that reality and act on it, and it is the natural first step in any answer engine optimisation effort.
What prompt testing actually is
Prompt testing is simply asking AI assistants the questions your customers ask, and recording how you show up. It treats the AI as your customer would, then reads the results like an audit rather than a chat, so what feels like idle experimenting becomes real data.
The reason it works is that AI answers are shaped by the prompt. As the practice of prompt engineering shows, small changes in wording can change what an answer contains, so testing a range of realistic prompts reveals a fuller picture than any single question.
You are not looking for one perfect answer, you are mapping the pattern across many.
It also respects how these systems really behave. OpenAI's own overview of ChatGPT's capabilities makes clear that it can search the web and that its responses vary, so a good test accounts for that variation rather than trusting a lone result. Done well, prompt testing turns a fuzzy worry about AI visibility into concrete, recorded evidence.
Step one: build your prompt set
A good audit starts with a good prompt list, drawn from how your customers actually think. Aim for fifteen to thirty prompts across a few types, written in plain, natural language rather than marketing terms.
Start with brand prompts that name you directly, such as "What do you know about [your brand]?" and "Is [your brand] any good?" Add category prompts that do not name you, the ones buyers really use, like "Who are the best [your service] in [your area]?" and "Recommend a [your category] for [a common need]."
Then include comparison prompts such as "[your brand] versus [a competitor]" and problem-based prompts like "How do I solve [the problem you fix]?"
Write them the way a real person types, because that is what the model responds to. Nielsen Norman Group's research on real generative-AI user behaviour shows people phrase things naturally and conversationally, so match that. A prompt list built from genuine customer language is the single biggest factor in a useful audit.
The best source of prompts is not your imagination, it is your customers. Mine your sales calls, support tickets and enquiry emails for the exact questions people ask before they buy, and turn each into a prompt. Note the words they use for your service, the comparisons they raise, and the worries they voice.
A list drawn from real conversations tests the answers your actual buyers will see, whereas a list you invented at your desk tends to flatter you by asking the questions you wish they asked. Spend the extra half hour building the list from evidence, because everything downstream depends on it.
Step two: run the prompts systematically
How you run the prompts matters as much as the prompts themselves, because AI answers vary. A little discipline here is what separates a reliable audit from a misleading one.
Run each prompt across the assistants your customers use, not just one, since a brand strong in one can be absent in another. Give Google's AI answers equal weight alongside the chat-first tools, because they reach an enormous audience and often draw on different signals.
Use a fresh chat for each prompt so previous questions do not colour the answer, and run the most important prompts a couple of times, because responses shift between sessions. Note whether live search was used, as that changes what the model can find.
Keep the conditions consistent so your results are comparable over time. Test from a location that matches your audience, use the same account settings each round, and record the date, since answers change as models update. This consistency is what makes a repeat audit next quarter genuinely comparable, rather than noise.
Step three: record the results
An audit is only as good as its record, so capture each result in a simple spreadsheet. One row per prompt, with columns for what you need to see at a glance.
Record whether you appeared at all, how prominently you were mentioned, whether the description was accurate, and who else appeared alongside or instead of you.
A quick yes or no for appearance, a note on prominence, a flag for any errors, and a list of competitors named is enough. Search Engine Journal's guidance on tracking AI visibility and prompts the right way reinforces that a consistent, recorded method beats ad-hoc checking, because it lets you see change rather than just a moment.
Capture the competitor detail especially, because it is gold. The brands that keep appearing when you do not are your benchmark, and studying how they are described tells you what a winning presence looks like in your category.
Keep the sheet simple enough that you will actually fill it in. A handful of columns you complete every time beats an elaborate template you abandon after the first round. The point is not a beautiful spreadsheet, it is a consistent record you can compare month to month, so favour speed and repeatability over detail.
If a column takes real effort to judge, either define it tightly so anyone would score it the same way, or drop it. The audit only works if you keep doing it, and simplicity is what keeps you doing it.
Step four: score and read the results
With the sheet filled in, turn it into something you can act on and track. The simplest useful number is a visibility rate: the share of your prompts where you appeared at all.
Break it down by prompt type, because the pattern is the insight. Appearing for brand prompts but not category prompts means you are known but not competitive. Appearing with wrong details means an entity or accuracy problem.
Never appearing means an invisibility problem to fix from the ground up, which our guide to being cited as a source helps you address. Add a note on sentiment and accuracy, since being mentioned badly is its own issue.
Then tie the findings to outcomes. The prompts closest to a purchase matter most, so weight your reading towards the category and comparison prompts that drive decisions, and connect the audit to measuring AEO return so visibility work stays tied to leads.
Resist the temptation to average everything into one tidy figure. A single blended score can hide the fact that you are winning the questions that do not matter and losing the ones that do.
Two numbers are usually enough: how visible you are overall, and how visible you are on the handful of high-intent prompts that actually produce customers. The second is the one to watch, because it is the number most closely tied to revenue.
Step five: when to scale with tools
Manual testing is perfect to start, but it has limits: you can only run so many prompts, so often, by hand. When you want more coverage or a continuous view, tools take over the heavy lifting.
Several are free to begin with. Semrush offers an AI search visibility checker that tracks your presence across assistants automatically, turning your one-off manual audit into an ongoing measurement.
Use tools to scale the frequency and breadth, while keeping the occasional manual run to see exactly how you are described, which a score alone never shows.
The right moment to add a tool is when manual testing has proven there is something worth tracking. Start free, learn what matters, then let a tool watch it continuously so you catch changes early.
Common mistakes to avoid
A few errors quietly ruin a DIY audit, so watch for them. Most come from treating a variable system as if it were fixed.
Do not trust a single run, since answers vary between sessions, and do not test only your brand name, which flatters you while hiding the category gaps that matter.
Avoid leading or marketing-style prompts that no real customer would type, because they produce answers no real customer will see. And do not test only ChatGPT, since your buyers spread across several assistants.
Fixing your presence then means optimising for ChatGPT search, Perplexity and Google's answers individually.
One more mistake is treating the test as a one-time verdict rather than a baseline. A single audit tells you where you stand today, but its real power is comparison, so its first run matters most as the reference point for every run after.
Resist the urge to overreact to one alarming answer or celebrate one flattering one.
Record it, move on, and let the pattern across many prompts and several rounds tell the real story, because that pattern is far more reliable than any single response a variable system happens to give you.
Make it a repeatable habit
A single audit is a snapshot, and AI moves. Models retrain, live search shifts, and your own presence changes, so the real value comes from running the same test regularly and watching the trend.
Keep your prompt list and spreadsheet, and re-run the audit on a set cadence, monthly or quarterly, logging your visibility rate each time.
Start from your home base and let a structured AI visibility audit extend your manual method across more prompts and platforms when you are ready.
Tie it to your generative engine optimisation plan, and our case study shows how tracking and acting on this kind of data recovers visibility over time.
Then act on what you find. A test that never changes anything is just curiosity, so use each round to fix the weakest results, whether that is entity clarity, source presence or content, guided by work like entity optimisation for AI and building brand trust in the AI era.
The audit shows the problem and the action closes it, and the same measure-then-fix loop underpins turning lost discovery into AI citations. Each cycle should end with a change, not just a chart.