The Clinical-Grade Question
Three things happened in the last three weeks that, taken together, expose a tension at the center of healthcare AI in 2026. CMS proposed the first-ever Medicare payment category for AI diagnostic software. Forty-three states are advancing more than 240 AI healthcare bills, with the first wave of laws now in effect. And a landmark study in Nature Medicine showed that the general-purpose AI tools clinicians can access for free already outperform the specialized, purpose-built clinical AI products that hospitals and vendors market as the professional-grade alternative. The question of what makes an AI tool "clinical-grade" is being answered simultaneously by regulators, researchers, and the market.
In today's newsletter:
CMS proposes "Software as a Medical Service," the first structured Medicare payment pathway for AI diagnostic tools, covering 36 procedure codes across radiology, cardiology, ophthalmology, and pathology
Indiana and Tennessee become the first states to restrict AI in clinical billing and mental health.
Neko Health raises $700 million to bring AI full-body scanning to the United States, backed by Mark Zuckerberg, Priscilla Chan, and Lightspeed Venture Partners
Nature Medicine study: general-purpose LLMs outperform FDA-cleared specialized clinical AI on real clinical queries
The Stanford-Harvard State of Clinical AI 2026 report on what actually works at the bedside versus what performs well in a lab
In most of medicine, "clinical-grade" has a reasonably clear meaning. A clinical-grade lab reagent meets validated purity and accuracy thresholds. A clinical-grade device has cleared a defined regulatory pathway. The term implies independent verification against a standard that the product had to meet before you were expected to trust it with patient care. For AI, that clarity does not yet exist, and this week's news makes the gap harder to ignore.
So many applications out there today, claim to be clinical grade, but are they really? Many models claim to outperform clinicians, but don’t usually hold water in real-world applications.
The answer to that question requires something none of the current credentialing systems require: prospective outcomes data from real clinical environments, at scale, over time. Clinical AI has boomed, and the strongest results are concentrated in prediction tasks in controlled settings. What breaks down when those tools leave the lab is the ability to handle uncertainty, and the risk that clinicians follow incorrect model recommendations even when the errors are detectable. The clinical-grade question is not just about who approves a tool or who pays for it. It is about whether the tool actually changes what happens to patients. Right now, for most AI in healthcare, we do not have a rigorous answer.
LATEST NEWS

From Stat News
CMS Proposes "Software as a Medical Service" — The First Medicare Payment Pathway for AI Diagnostics
CMS's 2027 OPPS proposed rule, released July 2, creates a new payment category called Software as a Medical Service (SaMS) and designates 36 HCPCS codes covering AI diagnostic tools — including AI analysis of retinal images, echocardiogram-based heart failure detection, CT-derived coronary blood flow estimates, EKG-based cardiac risk scoring, and AI-assisted prostate cancer mapping. CMS proposes a new "O1" status indicator and moves 21 of those codes into New Technology APCs, the temporary payment track it uses when claims data is insufficient to set a permanent rate. The agency calls the framework an interim policy and acknowledges it will need to develop a long-term valuation methodology as evidence accumulates.
What it means: For the first time, AI diagnostic software will have its own Medicare billing infrastructure — which changes the economics of AI adoption for hospital outpatient departments and ASCs. Health systems that have deferred AI diagnostic tool purchases should understand what this payment pathway covers and what validation evidence CMS will eventually require to make rates permanent.
State AI Healthcare Laws Take Effect July 1 — 43 States, 240+ Bills, No Federal Standard
Indiana and Tennessee became the first states to enact AI-specific healthcare restrictions effective July 1, 2026. Indiana's HB 1271 prohibits health insurers from using AI as the sole basis for downcoding a claim without clinician review, and prohibits providers from submitting AI-generated claims without human verification. Tennessee's SB 1580 prohibits AI chatbots from representing themselves as qualified mental or behavioral health professionals. The laws reflect two of the most active legislative themes in 2026: AI oversight in payer decisions, and patient-facing AI transparency in behavioral health. Forty-three states have now introduced more than 240 AI healthcare bills this year, nearly matching the total from all of 2025.
What it means: Health systems operating across multiple states now face a compliance patchwork with no federal baseline beneath it. The same AI documentation or prior authorization tool may be regulated differently — or not at all — depending on where the patient is seen. Legal and compliance teams need a state-by-state AI policy audit before the next wave of laws takes effect.
Neko Health Raises $700 Million to Bring AI Full-Body Scanning to the United States
Neko Health, the AI-powered body scanning company founded by Spotify's Daniel Ek, raised $700 million in a Series C round on July 15, valuing the company at $7 billion. The round was led by Lightspeed Venture Partners and included investment from Mark Zuckerberg, Priscilla Chan, and a roster of high-profile individual investors. Neko's 60-minute, non-invasive, radiation-free scan uses AI to assess cardiovascular, metabolic, and musculoskeletal health from a single appointment. The company has been operating in Europe and is using the capital to open its first US facilities. It now has more than $1 billion in total disclosed funding since 2023.
What it means: Neko is a direct-to-consumer model — patients book and pay directly, outside of insurance. At scale, that creates a two-tiered preventive care landscape: AI-powered full-body assessment for those who can afford the out-of-pocket cost, and standard preventive care for everyone else. Clinicians should be prepared for patients arriving with Neko scan results that fall outside standard diagnostic categories.
RESEARCH

General-Purpose Chatbots Outperform Clinical AI Tools on Physicians' Real-World Questions
Nature Medicine, June 2026
Researchers at Stanford and collaborating institutions evaluated three frontier general-purpose LLMs (GPT-5.2, Gemini 3.1 Pro, and Claude Opus 4.6) against two purpose-built clinical AI products (OpenEvidence and UpToDate Expert AI) across three benchmarks: 500 MedQA knowledge questions, 500 HealthBench alignment items, and a real clinical query benchmark built from 100 de-identified questions submitted by practicing physicians in a live clinical setting. The RCQ benchmark was reviewed by 12 US clinicians in randomized, blinded assessments. Frontier LLMs outperformed clinical AI tools on all three benchmarks. On real clinical queries, the specialized tools performed comparably to Google's AI search overview. Clinicians preferred the frontier LLM outputs in blinded review. Read the study.
Key Finding: General-purpose frontier AI models consistently outperform specialized clinical AI products on medical knowledge tasks, clinical alignment measures, and real physician queries — and practicing clinicians prefer them.
Clinical Implication: The study's authors are careful to note that benchmark performance does not capture EHR integration, regulatory compliance, or liability frameworks — all of which matter in practice. But the finding challenges the premise that a "clinical-grade" designation implies superior performance. It also raises a direct question for health systems paying for specialized clinical AI subscriptions: what exactly are they paying for?
State of Clinical AI 2026: What Holds Up in Practice
ARISE Network (Stanford-Harvard), January 2026
The inaugural State of Clinical AI report from the ARISE network — a Stanford-Harvard research consortium — synthesized the most significant evidence on clinical AI deployment across 2025. The report found that the strongest, most consistent results appear in prediction tasks, where AI analyzes large datasets to generate risk scores and early warning signals, and in radiology, where AI functions as an optional second opinion for image interpretation. Where performance breaks down: models struggle to identify their own uncertainty, and multiple studies documented over-reliance, with clinicians following incorrect AI recommendations even when errors were detectable by the clinician alone. The report draws an explicit distinction between controlled study performance and real-world deployment outcomes. Read the report.
Key Finding: Clinical AI performs best as a risk stratification and early warning tool. The biggest safety risk is not underperformance — it is clinician over-reliance on outputs the model itself should have flagged as uncertain.
Clinical Implication: Before adopting any AI clinical tool, ask specifically: how does this system communicate uncertainty? If the answer is that it does not — that it always returns a confident output — that is a red flag, regardless of benchmark accuracy. A tool that does not know when it does not know is more dangerous than one that occasionally gets the right answer wrong.
ETHICS/REGULATION

From Manatt Health
43 States Are Writing Their Own AI Healthcare Rules — and the Definitions Are Not Compatible
With no comprehensive federal AI healthcare legislation in effect, states have become the primary regulatory force shaping how AI can be used in clinical care, insurance decisions, and patient-facing health tools. Indiana prohibits AI-only claim downcoding. Tennessee prohibits AI chatbots from posing as mental health professionals. Colorado requires disclosure when AI is used in high-stakes decisions. Other states are debating consent requirements, liability assignment, and mandatory human review thresholds — often using different definitions of what counts as AI, what counts as clinical use, and what counts as high-stakes. A tool that is compliant in one state may require modification, additional disclosure, or clinical oversight workflows in the next state over. Manatt's Health AI Policy Tracker is currently the most comprehensive resource for following this landscape in real time.
Why This Matters: Health systems and health technology vendors that operate across state lines are now building compliance infrastructure for a moving, fragmented target. Clinicians practicing in multi-state systems should ask their legal and compliance teams for a current-state AI policy map — and should assume it will need to be updated before the end of the year.
FINAL THOUGHTS
The question of what makes an AI tool "clinical-grade" is not going to be resolved cleanly or soon. CMS is building a payment pathway with interim rates and no permanent methodology. States are writing laws with definitions that do not align with each other or with the FDA's. And the benchmark evidence suggests that the designation itself may not reliably predict which tools perform better in practice. That is not a reason for despair. It is a reason for precision.
Clinicians have always had to evaluate tools in the absence of perfect information. We adopted statins before we fully understood pleiotropic effects. We deployed laparoscopic surgery before long-term outcome data existed at scale. The pattern of adopting promising tools while evidence accumulates is not new to medicine. What is new is the speed of deployment, the opacity of the underlying models, and the fact that failure modes in AI often do not look like failure — they look like confident, plausible, wrong answers. That is a different category of risk than a drug with a known side effect profile.
What responsible adoption looks like in this environment is not waiting for regulatory certainty before using any AI tool. It is being specific about which clinical decisions you are and are not asking AI to influence, knowing enough about how a tool was built and validated to judge its limitations, and maintaining the clinical independence to override it when your judgment and the model's output do not agree. The credential matters less than the judgment you bring to reading it.
Start Here:
Check out this week’s clinical Guide: How to Evaluate AI Tools in a Fragmented Regulatory Landscape
Best Regards,
Chris Massey, MD
"All models are wrong, but some are useful."
Are you enjoying Intelligent Medicine?
Send me an email letting me know what you’d like me to discuss in future issues @ [email protected]
Disclaimer: This newsletter is for educational and informational purposes only and does not constitute medical advice. Readers should review primary sources and follow applicable clinical guidelines and institutional policies before implementing any changes. Always de-identify patient data and review all outputs for accuracy.
