Medicine has carried a billing taxonomy for AI since 2022. This spring it was quietly tightened — and the revision reads less like a payment rule than a clinical thesis about judgment.
The story most people will tell about the CPT Editorial Panel's May meeting is that medicine made it easier to bill for artificial intelligence. That is not quite what happened. The Panel revised Appendix S — the code set's taxonomy for AI-enabled services — and the effect was to make the bar higher, not lower. Before we can read why that matters, we have to be clear about what this appendix is, because nearly every conversation about it starts by getting that wrong.
Appendix S does not classify your use of a chatbot to pick a code. It classifies AI-enabled medical services and procedures — the work a machine performs on behalf of the physician — into three tiers, each defined by how much clinical judgment the machine displaces. It exists for one reason: to give AI-powered procedures a defensible path to payment.
The Panel did not add a fourth tier. It raised the evidentiary discipline inside the three. Assistive was narrowed: outputs hedged as "risk for" or "suggestive of" may now demand clinical validation before they qualify. Augmentative was raised: an output has to be clinically meaningful and pertinent to the code descriptor — not merely statistical. Autonomous was tightened around transparency, guidelines, and demonstrated clinical utility. A loose vocabulary became an evidence standard. To claim a tier now, you have to prove the machine earned it.
Read structurally, this is the payment system encoding the one habit every clinician is supposed to bring to a machine's output. Two of the three tiers legally require a physician to interpret and report — the human is not a courtesy, it is the billable act. And the revision now forces a builder to demonstrate clinical meaning before climbing to a higher tier. The code set, in other words, refuses to take the output on faith.
Here is the line worth holding, because blurring it is the most common error in this debate. Appendix S is about AI as a billable service — the machine does diagnostic work that earns a code. It says nothing about AI as a coding assistant — an LLM assigning your E/M level from a note. That second frontier is a different question entirely, and it is the one I study directly: whether prompt structure changes how accurately a model codes against a human gold standard. Treating the two as one story is how smart people end up confidently wrong. They are adjacent. They are not the same.
The taxonomy exists, the payment pathway exists, and foot-and-ankle AI is barely represented in either. That is not a deficit to mourn — it is runway. Podiatry is small enough to move deliberately and close enough to the patient to set its own evidence bar before the codes arrive, rather than inheriting someone else's after the fact.
For the resident or student reading this: learn this framework now. It is how the AI in your future practice will — or won't — get paid for, and it models the posture the rest of medicine is only now writing into policy. Understand what the machine is allowed to conclude before you trust what it does.