HomeInsights
Elevate

Prompt Testing: The DIY Way to Audit Your AI Visibility

You do not need an agency or a paid tool to find out how you look in AI answers. Here is a simple, repeatable way to audit your own AI visibility with nothing but a list of prompts and a spreadsheet.

Man in white shirt and patterned vest sitting in a black leather chair at a table, resting chin on fist.

Eden John

Founder, SkyScale

6 min read

Published

July 24, 2026

Updated

July 24, 2026

Decorative

What changed in this article (July 24, 2026): Rebuilt as a practical, do-it-yourself prompt-testing method using OpenAI documentation on how ChatGPT answers, Nielsen Norman Group research on real AI behaviour, and Search Engine Journal guidance on tracking prompts the right way.

Table Of Content

Quick summary

Prompt testing is the free, do-it-yourself way to see exactly how AI answers describe and recommend you. This guide gives you a repeatable method to build, run, record and read your own AI visibility audit.

  • Prompt testing needs only a prompt list and a spreadsheet.
  • Test brand, category and comparison prompts, not just your name.
  • Run each across assistants, in fresh chats, more than once.
  • Record whether you appear, how, and who beats you.
  • Score it into a visibility rate you can track over time.
Audience Icon

Who this is for

This is for owners and marketers who want to test their AI visibility themselves, before paying anyone, and get real data rather than guesswork.

  • Owners who want a free baseline before hiring help.
  • Marketers building a repeatable AI visibility check in-house.
Evidence base document icon

Evidence base

This draws on OpenAI documentation on how ChatGPT answers, Nielsen Norman Group research on how people use generative AI, Search Engine Journal guidance on tracking AI visibility and prompts, Pew Research data on AI adoption, and patterns from more than 200 AI visibility audits we ran between October 2024 and June 2026.

Research methodology icon

Methodology

We distilled the manual prompt-testing process we use at the start of every audit into a repeatable method anyone can run with free tools.

Limitations warning icon

Limitations

AI answers vary by phrasing, session, location and model version, so no single run is definitive. This is a directional method, not a precise measurement, and vendor supplied performance claims were excluded in favour of official and independent sources.

Desk arranged with prompt cards, arrows, check marks and a checklist, representing a DIY audit of AI visibility through prompt testing.

Why do it yourself first

Before spending on tools or agencies, there is real value in testing your AI visibility by hand. It is free, it takes an hour, and it gives you something no dashboard can: the actual words AI uses to describe you, in the answers your customers see.

Doing it yourself also builds intuition. When you run the prompts personally and watch what appears, you develop a feel for how AI answers your category, who it favours, and where you stand. That understanding makes every later decision sharper, whether you eventually buy a tool or not, and it costs nothing but attention.

There is a confidence benefit too. It is easy to feel anxious and powerless about AI visibility when it is an abstract worry, but running the test turns that feeling into information you can act on.

You may discover you are in better shape than you feared, or you may confirm a real gap, but either way you are no longer guessing. That shift from vague dread to concrete findings is worth the hour on its own, and it is why we start every engagement with exactly this kind of manual look before recommending anything.

There is a reason to start now rather than later. With AI adoption rising fast, as Pew Research documents in its work on how people are using AI in daily life, the answers being formed about you today are already shaping decisions. A manual audit is the fastest way to see that reality and act on it, and it is the natural first step in any answer engine optimisation effort.

What prompt testing actually is

Prompt testing is simply asking AI assistants the questions your customers ask, and recording how you show up. It treats the AI as your customer would, then reads the results like an audit rather than a chat, so what feels like idle experimenting becomes real data.

The reason it works is that AI answers are shaped by the prompt. As the practice of prompt engineering shows, small changes in wording can change what an answer contains, so testing a range of realistic prompts reveals a fuller picture than any single question.

You are not looking for one perfect answer, you are mapping the pattern across many.

It also respects how these systems really behave. OpenAI's own overview of ChatGPT's capabilities makes clear that it can search the web and that its responses vary, so a good test accounts for that variation rather than trusting a lone result. Done well, prompt testing turns a fuzzy worry about AI visibility into concrete, recorded evidence.

Step one: build your prompt set

A good audit starts with a good prompt list, drawn from how your customers actually think. Aim for fifteen to thirty prompts across a few types, written in plain, natural language rather than marketing terms.

Start with brand prompts that name you directly, such as "What do you know about [your brand]?" and "Is [your brand] any good?" Add category prompts that do not name you, the ones buyers really use, like "Who are the best [your service] in [your area]?" and "Recommend a [your category] for [a common need]."

Then include comparison prompts such as "[your brand] versus [a competitor]" and problem-based prompts like "How do I solve [the problem you fix]?"

Write them the way a real person types, because that is what the model responds to. Nielsen Norman Group's research on real generative-AI user behaviour shows people phrase things naturally and conversationally, so match that. A prompt list built from genuine customer language is the single biggest factor in a useful audit.

The best source of prompts is not your imagination, it is your customers. Mine your sales calls, support tickets and enquiry emails for the exact questions people ask before they buy, and turn each into a prompt. Note the words they use for your service, the comparisons they raise, and the worries they voice.

A list drawn from real conversations tests the answers your actual buyers will see, whereas a list you invented at your desk tends to flatter you by asking the questions you wish they asked. Spend the extra half hour building the list from evidence, because everything downstream depends on it.

Step two: run the prompts systematically

How you run the prompts matters as much as the prompts themselves, because AI answers vary. A little discipline here is what separates a reliable audit from a misleading one.

Run each prompt across the assistants your customers use, not just one, since a brand strong in one can be absent in another. Give Google's AI answers equal weight alongside the chat-first tools, because they reach an enormous audience and often draw on different signals.

Use a fresh chat for each prompt so previous questions do not colour the answer, and run the most important prompts a couple of times, because responses shift between sessions. Note whether live search was used, as that changes what the model can find.

Keep the conditions consistent so your results are comparable over time. Test from a location that matches your audience, use the same account settings each round, and record the date, since answers change as models update. This consistency is what makes a repeat audit next quarter genuinely comparable, rather than noise.

Step three: record the results

An audit is only as good as its record, so capture each result in a simple spreadsheet. One row per prompt, with columns for what you need to see at a glance.

Record whether you appeared at all, how prominently you were mentioned, whether the description was accurate, and who else appeared alongside or instead of you.

A quick yes or no for appearance, a note on prominence, a flag for any errors, and a list of competitors named is enough. Search Engine Journal's guidance on tracking AI visibility and prompts the right way reinforces that a consistent, recorded method beats ad-hoc checking, because it lets you see change rather than just a moment.

Capture the competitor detail especially, because it is gold. The brands that keep appearing when you do not are your benchmark, and studying how they are described tells you what a winning presence looks like in your category.

Keep the sheet simple enough that you will actually fill it in. A handful of columns you complete every time beats an elaborate template you abandon after the first round. The point is not a beautiful spreadsheet, it is a consistent record you can compare month to month, so favour speed and repeatability over detail.

If a column takes real effort to judge, either define it tightly so anyone would score it the same way, or drop it. The audit only works if you keep doing it, and simplicity is what keeps you doing it.

Step four: score and read the results

With the sheet filled in, turn it into something you can act on and track. The simplest useful number is a visibility rate: the share of your prompts where you appeared at all.

Break it down by prompt type, because the pattern is the insight. Appearing for brand prompts but not category prompts means you are known but not competitive. Appearing with wrong details means an entity or accuracy problem.

Never appearing means an invisibility problem to fix from the ground up, which our guide to being cited as a source helps you address. Add a note on sentiment and accuracy, since being mentioned badly is its own issue.

Then tie the findings to outcomes. The prompts closest to a purchase matter most, so weight your reading towards the category and comparison prompts that drive decisions, and connect the audit to measuring AEO return so visibility work stays tied to leads.

Resist the temptation to average everything into one tidy figure. A single blended score can hide the fact that you are winning the questions that do not matter and losing the ones that do.

Two numbers are usually enough: how visible you are overall, and how visible you are on the handful of high-intent prompts that actually produce customers. The second is the one to watch, because it is the number most closely tied to revenue.

Step five: when to scale with tools

Manual testing is perfect to start, but it has limits: you can only run so many prompts, so often, by hand. When you want more coverage or a continuous view, tools take over the heavy lifting.

Several are free to begin with. Semrush offers an AI search visibility checker that tracks your presence across assistants automatically, turning your one-off manual audit into an ongoing measurement.

Use tools to scale the frequency and breadth, while keeping the occasional manual run to see exactly how you are described, which a score alone never shows.

The right moment to add a tool is when manual testing has proven there is something worth tracking. Start free, learn what matters, then let a tool watch it continuously so you catch changes early.

Common mistakes to avoid

A few errors quietly ruin a DIY audit, so watch for them. Most come from treating a variable system as if it were fixed.

Do not trust a single run, since answers vary between sessions, and do not test only your brand name, which flatters you while hiding the category gaps that matter.

Avoid leading or marketing-style prompts that no real customer would type, because they produce answers no real customer will see. And do not test only ChatGPT, since your buyers spread across several assistants.

Fixing your presence then means optimising for ChatGPT search, Perplexity and Google's answers individually.

One more mistake is treating the test as a one-time verdict rather than a baseline. A single audit tells you where you stand today, but its real power is comparison, so its first run matters most as the reference point for every run after.

Resist the urge to overreact to one alarming answer or celebrate one flattering one.

Record it, move on, and let the pattern across many prompts and several rounds tell the real story, because that pattern is far more reliable than any single response a variable system happens to give you.

Make it a repeatable habit

A single audit is a snapshot, and AI moves. Models retrain, live search shifts, and your own presence changes, so the real value comes from running the same test regularly and watching the trend.

Keep your prompt list and spreadsheet, and re-run the audit on a set cadence, monthly or quarterly, logging your visibility rate each time.

Start from your home base and let a structured AI visibility audit extend your manual method across more prompts and platforms when you are ready.

Tie it to your generative engine optimisation plan, and our case study shows how tracking and acting on this kind of data recovers visibility over time.

Then act on what you find. A test that never changes anything is just curiosity, so use each round to fix the weakest results, whether that is entity clarity, source presence or content, guided by work like entity optimisation for AI and building brand trust in the AI era.

The audit shows the problem and the action closes it, and the same measure-then-fix loop underpins turning lost discovery into AI citations. Each cycle should end with a change, not just a chart.

Implementation checklist

Use this list to audit and improve your AI visibility after reading this guide.

  • Build a list of fifteen to thirty prompts in real customer language.
  • Cover brand, category, comparison and problem-based prompts.
  • Run each across several assistants, in fresh chats, more than once.
  • Keep conditions consistent and record the date each round.
  • Log appearance, prominence, accuracy and competitors per prompt.
  • Score a visibility rate and break it down by prompt type.
  • Add a free tool once manual testing proves what is worth tracking.
  • Re-run on a set cadence and fix the weakest results each time.

Sources and references

Primary sources, official documentation, research and SkyScale audit data cited in this article. in this article.

Frequently Asked

What is prompt testing for AI visibility?

Decorative

It is the practice of asking AI assistants the questions your customers ask and recording how your brand shows up. You run a set of realistic prompts, note whether you appear, how you are described and who else appears, then read the pattern like an audit. It is the simplest way to see your real AI presence.

Do I need to pay for a tool to test my AI visibility?

Decorative

No. You can run a thorough audit for free using ChatGPT and a spreadsheet. Paid tools help you scale to more prompts and track continuously, but the manual method gives you the actual wording of the answers and a solid baseline. Start free, and add a tool only once you know what is worth tracking.

How many prompts should I test?

Decorative

Aim for fifteen to thirty, spread across brand, category, comparison and problem-based questions. That range is enough to reveal a clear pattern without becoming a chore. The mix matters more than the number: category prompts that do not name you are the most revealing, because they show whether you surface when buyers are actually deciding.

Why do I get different answers each time I test?

Decorative

Because AI answers vary by phrasing, session, location and whether live search runs. That is normal, not a fault. It is exactly why you run key prompts more than once, in fresh chats, and record the conditions. Reading the pattern across several runs gives you a truer picture than any single response ever could.

How do I turn my results into a score?

Decorative

Use a visibility rate: the share of prompts where you appeared at all, then break it down by prompt type. Add notes on prominence, accuracy and sentiment. Tracking that rate over time, rather than obsessing over any one answer, tells you whether your AI presence is improving and which prompt types still need work.

How often should I run a prompt audit?

Decorative

Monthly or quarterly suits most businesses, because AI knowledge and answers change as models update and your presence grows. Keep your prompt list and spreadsheet so each round is comparable, log your visibility rate every time, and use the results to fix the weakest areas. Consistency turns testing into an early-warning system.

Authorship and review

Man in white shirt and patterned vest sitting in a black leather chair at a table, resting chin on fist.

Written by

Eden John

· Founder, SkyScale

 LinkedIn profile

Eden leads SkyScale's Generative Engine Optimisation practice, focused on getting brands cited inside ChatGPT, Perplexity, Google AI Overviews and Gemini.

Relevant experience: Shipped 100+ AI visibility audits across B2B SaaS, professional services and ecommerce between Q4 2024 and Q1 2026, tracking citation patterns across the four major answer engines.

Credentials: Master of Business Administration (MBA) · Founder, SkyScale · 100+ AI visibility audits · GEO, AEO and AI SEO specialist

Smiling young man with curly dark hair in a maroon T-shirt crosses his arms indoors.

Reviewed by

Lachlan McDonald

· AI Search & Data Engineering Reviewer

 LinkedIn profile

Lachlan reviews SkyScale's AI search and data engineering content, focused on technical accuracy, methodology, retrieval logic, data quality and source-evaluation claims.

Relevant experience: 6 years of experience across AI search and data engineering, reviewing technical systems and source-selection claims for accuracy, reliability and methodological soundness.

Credentials: Master of Data Science · Bachelor of Software Engineering (Honours) · AI search and data engineering specialist

Last reviewed March 27, 2026
This is the block containing the Collection list that will be used to generate the "Previous" and "Next" content. You can hide this block if you want.
Ai visibility icon

AI Visibility
Report

3 business days. No credit card required, reviewed by a human.

Real Client Results

What you can expect to gain

+1,975%

more clicks from search

Benarrivati

£2,262

revenue from ChatGPT

Avenue Cookery

Google CTR lift

Vision One

+462%

more search impressions

SkyScale

See how we did it