AEO tracking comes down to one repeated question: when a patient asks an AI assistant for a practice like yours, does the answer say your name? Picture the practice manager at a three-doctor ENT clinic asking her phone which ENT near her is best for chronic sinus problems, and getting back a tidy paragraph that names four practices without hers. She asks again a few weeks later, gets a different answer that still leaves her out, and has no way to tell whether anything in between moved. Closing that gap is the job: you check on a fixed schedule, you log what you see, and a hunch becomes a trend line. What’s really working in 2026 is slightly different from what most clinics set up first, and we will get there by the end.

One note on numbers: search volumes, difficulty scores, and cost-per-click figures in this article come from our own keyword research, a pull of 1,341 healthcare keywords, July 2026. Outside facts are linked to their primary source. Hours, channel rankings, and timelines are LabRanked planning estimates, not measured benchmarks.

Why this matters right now

Patients are already asking. About a third of the public, 32%, reports turning to AI chatbots for physical or mental health advice in the past year, per the KFF Tracking Poll on Health Information and Trust. In the same poll, 27% used AI to look up symptoms or conditions and 16% used it to help decide whether to see a doctor at all. Which practice gets picked after that is not something KFF measured, and we could not find an allowed source that did, so we are not printing a number for it. The first pass at a health question now happens somewhere your name either appears or does not.

The measurement problem arrives with the traffic. Pew Research Center’s metered study of 900 U.S. adults found that when an AI summary appeared, people clicked a traditional search result in 8% of visits, against 15% with no summary, and clicked a link inside the summary itself in only 1% of visits. Pew measured clicks rather than clinic visibility, but the consequence is not hard to see: held against those numbers, a good month can read like a bad one, because the appearance that mattered never produced a session.

Space inside the answer is tighter than a page of links, too. In the same study, 88% of summaries cited three or more sources, only 1% cited a single one, and the median summary ran 67 words. That counts linked sources rather than named businesses, and it is the shape of the surface: a short paragraph carrying a handful of citations.

The difference from classic search is who else is paying attention. Our own keyword data has “aeo tracking” drawing 210 searches a month at a difficulty score of 1 out of 100, which is what a discipline looks like before the crowd arrives, and most of the clinics we audit are not watching it yet. That is the part with a shelf life.

What used to work, and what it’s still good for

The measurement habits in a clinic’s marketing folder were built for a page of links. None of them are junk, but they are pointed at the wrong surface now, and if they stay the whole scoreboard they will hide the thing you are trying to see.

Rank tracking a keyword list. Still the right instrument for classic organic results, and worth keeping for that. It has nothing to report about an AI answer, because there is no position four inside a paragraph. You are named or you are not, and the paragraph gets rewritten for whoever asks next.

The monthly traffic report as the scoreboard. Sessions can fall for reasons that have nothing to do with how often you were the answer, which is what the click numbers above describe. Keep the report and demote it: it answers how many people arrived, and the question now is how often you were the answer.

The one good screenshot. Everybody has the one where an assistant named them, forwarded to the group chat. Answers vary by session, wording and location, so one reading proves it happened once, and twenty readings a month with the wording frozen is closer to a measurement.

Bolting AI files onto the site. An llms.txt file, content chopped into chunks, extra markup for the robots. Google’s AI optimization guide says you do not need new machine readable files, AI text files, markup, or Markdown to appear in Google Search including its generative AI capabilities, because Search does not use them, and it files seeking inauthentic mentions under the things owners can ignore.

Waiting on a single AI visibility score. Several products will sell you one. We could not find an independent evaluation of how any of them sample the engines, which makes the score a black box pointed at a moving target. If you buy one, keep your own log next to it: your log is the part you can audit.

The common thread is that these all measure a position on a page. This layer does not serve pages: it writes a fresh answer for one person, and your name is in it or it is not.

What’s working now, ranked

MethodCostSpeedVerdict for 2026
Search Console, Web Performance reportFreeAlready collectingWhere we start. Every property has it
Search Console, Generative AI reportFreeWhere you have accessSubset rollout, counts your URLs
A frozen prompt list, rerun monthlyFree, about 2 hours a monthSame dayThe core of the practice
Logging what each answer citedFree, rides on the prompt runSame dayWhere the mention has to come from
Conversions instead of sessionsFreeOngoingThe read we trust on AI traffic
The intake question at bookingNear zeroWeeksTies an answer to a booked patient
Paid AI visibility dashboardsPaidFastNo independent accuracy check we found

That ranking is our own framework rather than measured data. The split matters more than the order: Google’s reports tell you what happened to your URLs, the prompt runs tell you what the patient sees. Neither is the whole picture, which is why the setup below runs both.

Start where Google counts you. Google’s docs say sites appearing in AI features are included in overall search traffic in Search Console, in the Performance report under the Web search type, and that a page needs no extra technical work to be eligible beyond being indexed and eligible to show with a snippet. Every verified property has that number already, free and collecting. Google also runs a separate Generative AI performance report, and its help page says that one is rolling out to a subset of website owners, so not all properties have access. Where you have it, it shows impressions from generative AI features on Search by page, country and device. Read the word impressions carefully: it counts your URLs appearing, not the answer saying your name, and neither report tells you which practice got named in your place.

The prompt list. Write 20 questions the way a patient types them, longer and messier than the way a marketer would, then freeze the wording. Pew’s numbers help you choose: 53% of searches with ten or more words produced an AI summary against 8% of one and two word searches, and 60% of queries starting with a question word produced one. Those describe Google searches rather than every assistant, so use them for shape, not law. “Who treats vertigo near me in Fort Worth and takes Aetna insurance” is the long, question-shaped kind of prompt that tends to pull a generated answer. A bare “ENT” is not. Cover your top three services, the insurance question and the cost question. For every run, log whether you were named, where in the answer, what it said about you, and which sources it cited. That last column is the one people skip, and it says where the mention has to be earned.

Read what got cited. In that same Pew study, .gov sites were 6% of AI summary sources against 2% of standard results, and Wikipedia, YouTube and Reddit together made up 15%. That counts which domains got linked rather than what an engine trusts, but the lesson survives the distinction: answers are assembled out of pages that already exist somewhere, and a clinic that exists as a homepage and a phone number is thin material. In the AI visibility audits we run, we check first whether an assistant names the practice at all in its own market, and where it does not, the site usually has little specific enough to lift.

Give it something quotable. A peer-reviewed paper published in the proceedings of KDD ‘24 built a benchmark of 10,000 queries across 25 domains including health and reported its main results on the benchmark’s 1,000 query test split. Three separate edits, adding citations, quotations and statistics, each produced a relative improvement of 30 to 40% on a position-adjusted word count metric in the engines the authors tested, while keyword stuffing gave little to no improvement. Read it narrowly: those engines were a research setup and Perplexity rather than Google’s AI features, health was one of 25 domains rather than a separate result, and the study worked from a fixed set of sources, so it says nothing about what gets retrieved. What it supports is smaller and more useful. Once your page is in the pool an engine draws from, specific, attributed, checkable facts make it more visible inside the answer.

Stop grading yourself on clicks. Google’s Search Central blog says that clicks from results pages with AI Overviews are higher quality, with users more likely to spend more time on the site. Set that next to the drop in click rate and the scoreboard here is calls, forms and booked visits, with sessions as a footnote.

Ask at the front desk. Add one line to intake: how did you hear about us, with “an AI assistant” as something a person can tick. It is the one line in this list that ties a named answer to a patient in a chair. We could not find published data on how often patients volunteer that a chatbot sent them, so treat it as our habit rather than a finding, and expect a small number at first.

The 30-day action plan

  1. Days 1 to 2, open Google’s own numbers. Confirm Search Console is verified on the right property and read the Performance report under the Web search type. Check whether the generative AI view is on for you, and work from the Web report if not. About 1 hour.
  2. Days 1 to 3, write the prompt list. Twenty questions in patient wording with your city inside them, in one sheet nobody edits afterward. About 2 hours.
  3. Week 1, run the baseline. Every prompt, on every assistant your patients use, logged four ways: named, position, what it said, what it cited. About 3 hours.
  4. Week 2, fix what an engine cannot read. Profile accuracy first, then rewrite your top three service pages so the first two sentences answer the question and carry a specific fact. About 4 hours.
  5. Week 3, close the loop to revenue. Add the intake line and set the conversion view so calls and forms are the headline number. About 1 hour.
  6. Week 4, rerun the identical list and diff it. Same wording, same day, same order, noting what you published in between. About 2 hours.

Cash cost is close to nothing. Those hours are LabRanked planning estimates rather than measured benchmarks, and they total 13 across the month. Month two costs about two hours, which is the point of freezing the wording. In our experience, whether you are still running it in month five is what decides the outcome.

So what’s really working in 2026?

It is not a tool. It is a baseline.

Everything on the old list measured a position on a page that holds still. AI answers do not hold still, so the same question asked twice comes back different, and one reading is closer to weather than to climate. Which is why the boring version wins: the same twenty prompts, the same wording, the same week of the month, in a sheet nobody redesigns. Three months in, that sheet holds something we have not been able to buy, because it records what you published in between.

That is where the tracking earns its keep. Knowing which of your pages an answer pulled from, and which question it still cannot answer about you, tells you what to write on Monday. In the audits we run, the practices getting named tend to be the ones that handed an engine something exact enough to lift. The fundamentals underneath are the same unglamorous ones that feed the map pack and the organic results.

The tracking is just how you find out.