In Vivo

Observations on AI in healthcare

AI for Providers

Where are the AI doctors?

Ask the average physician how most of their day is spent and the answer is rarely “with patients.”

The more likely response is that it is spent wading through a sea of administrative tasks: documentation, prior authorizations, inbox messages, and the like.

A time-and-motion study in 2016 found physicians spend about two hours on the electronic health record and desk work for every hour of direct patient care, plus another 1-2 hours after-hours at home each night. A 2024 follow-up found the situation had barely improved nearly a decade later.

It makes some sense that this is the unglamorous site of AI’s first real foothold in care delivery.

This piece maps what is being built today, who is paying for it, and the economic pressure that explains why paperwork has been the first bastion of AI in healthcare. We’ll be intentionally light on a few topics that deserve fuller treatment (namely, AI in surgery).

Context

US hospitals have been running remarkably thin as of late.

The median hospital had operating margins of 0.2 percent in 2022, the sector’s worst year since the start of the pandemic, with roughly half of hospitals finishing in the red. Margins recovered to about 2 percent in 2023 and approached 5 percent in 2024, but even in recovery about 39 percent of hospitals lost money in 2023. In summary: a low-single-digit-margin business, with very little cushion.

The single largest line item is labor, about 56 percent of hospital expenses. The pandemic tested that further, exacerbating an already material nursing shortage. At the January 2022 peak, travel nurses were 23 percent of nurse hours but nearly 40 percent of nurse labor spend. Agency rates charged to hospitals rose more than 200 percent over three years. Contract-labor pressure has somewhat receded, but the memory remains close.

The nursing shortage is structural. About 100,000 RNs left the workforce during the pandemic. The median nurse is now 50 years old with roughly one in five over 65. Each RN who leaves costs a hospital about $60,000 to replace, an average of roughly $5 million a year per hospital. Federal projections put the national RN shortage at about 10 percent in 2027, worse in rural areas.

The situation of physicians is not terribly different: the AAMC projects a shortfall of up to 86,000 physicians by 2036, with about 20 percent of the current workforce already over 65. Demand is moving the other way, as the 65-and-older population heads toward about 71 million by 2030 and Medicare enrollment approaches 80 million.

Then there is the volume of the underlying labor - i.e., administrative load - itself. Administrative complexity is the single largest category of wasteful US health spending, on the order of $265 billion a year, and administration overall runs roughly a quarter to a third of national health spending. Prior authorization alone consumes about 13 physician-and-staff hours per practice each week, and 40 percent of physicians employ staff who work on it exclusively.

The highest-return, lowest-clinical-risk place to put AI first is documentation and administration. It helps that this is toilsome work, which nobody particularly likes, and that prior auths are not spiritually central to the identity of HCPs.

Mapping the field

There are (I’m sure) many taxonomies of AI in healthcare - here lies another:

First, augmentation or automation? Does a human remain in the loop and leverage output (augmentation), or does the system act autonomously (automation)?

We’ll see below the former is much more common today. There’s a good reason for this, beyond regulation and risk tolerance: it is a practical way to gain real-world data to train models.

Every scribe a physician corrects and every reading a radiologist confirms is a labeled example. A copilot that demands doctors fix its output is effectively capturing training data for reinforcement learning from human feedback (as Counsel Health’s CEO pointed out).

Second, proximity to clinical decisionmaking. How close does the tool sit to the moment a diagnosis is made or a treatment is chosen? Billing and scheduling: far. Notes and knowledge search: closer. Reading a scan or suggesting a diagnosis: closer still.

Practically, the two are heavily linked. The nearer a product is to a clinical decision, the riskier it is to let it act alone. But this isn’t necessarily the case, if performance is sufficient and incentives are aligned - more on this shortly.

Furthest from the patient and where automation runs freest. A wrong code is a financial error (and an impactful one at that), but typically not a clinical one.

  • Revenue cycle management. Commure, merged with Athelas at about a $7 billion valuation, says it handles more than 85 percent of revenue-cycle work with no human in the loop, across a base that includes roughly 188 HCA hospitals.
  • Autonomous coding. Nym Health, Fathom, and CodaMetrix turn clinical notes directly into billing codes. Nym markets “zero human intervention,” Fathom claims more than 90 percent end-to-end, and CodaMetrix, spun out of Mass General Brigham, codes the confident majority and routes the rest to a person.
  • Prior authorization. Cohere Health and Anterior automate approvals but send every denial to a clinician: speed up a “yes”, don’t touch a “no”.

Documentation

One step closer lies clinical documentation, perhaps the busiest category in healthcare AI. Ambient scribes listen to visits and draft notes. Roughly $5.6 billion flowed into the category in 2024.

A scribe is augmentative, since the clinician reviews and signs every note, and it is emphatically not a medical device (thereby avoiding any FDA clearance process). Low regulatory risk and a real pain point.

Also, where many of the biggest checks have gone:

  • Abridge. About $117 million in annual recurring revenue and a $5.3 billion valuation, deployed across Kaiser, Mayo, Johns Hopkins, and Duke.
  • Ambience Healthcare. A $243 million round at $1.25 billion, chosen by Cleveland Clinic after a five-vendor comparison.
  • Nabla and Suki. Other well-funded startups in the same lane.
  • The platforms. Everyone with distribution entered: Microsoft’s Dragon Copilot (600-plus organizations, more than 3 million conversations a month), Amazon’s HealthScribe, Oracle’s voice-first clinical agent, and Epic’s own native scribe.

(Because so much of this runs inside Epic, the fight for the ambient slot embedded in the dominant EHR is its own contest - one Epic, Microsoft, and Abridge are all in.)

Why scribes? Near-immediate payback.

A 2026 UCSF study of 1.2 million encounters found physicians who adopted a scribe billed about 1.8 more RVUs per week, roughly $3,000 per physician per year, with no rise in claim denials.

The authors could not discern whether that reflects more services, more accurate coding, or upcoding. A financial one-two punch: saving clinician time and improving reimbursement both raise top-line. By 2026, insurers and health systems broadly agreed that scribes increase billing intensity.

There is something to be said about the game theory this opens up, and how this will change plan design and insurer review (for another post).

Set against US health spending of $4.9 trillion (about 18 percent of GDP), so concentrated that top 5 percent of patients account for half, a productivity gain that shows up as heavier coding is not obviously a saving for the system. The starting point matters here, of course: is this a correction or drift?

Knowledge

Alongside documentation is the next generation of knowledge search products in medicine. The largest is OpenEvidence, a literature-grounded answer engine used by roughly 40 percent of US physicians and valued at $12 billion after a January 2026 round.

Like the scribes, OpenEvidence stays on the safe side of the line by surfacing evidence a clinician reviews, which qualifies as non-device clinical decision support. Navina, Regard, and Atropos Health sit close by (Regard seems to reach somewhat further, ingesting the EHR and suggesting diagnoses for review).

OpenEvidence’s business model is telling: it’s free to any physician with a valid NPI. The buyer is pharma, who buys prescriber attention at the moment of treatment choice (taking a page out of Google’s book). A pessimistic reading is that OpenEvidence has hit a $12 billion valuation as an advertising platform pointed at one of the most valuable moments in medicine.

It’s not clear how long the reign of OpenEvidence will last. By mid-2026, general-purpose frontier models (GPT-5.2, Gemini 3.1, Claude Opus 4.6) matched or beat purpose-built clinical tools like OpenEvidence and UpToDate’s AI on medical-question benchmarks. Those are benchmarks of course (a bit on that below), and in imaging - where the FDA has cleared autonomous tools - specialized tools still lead. But it does complicate the assumption that dedicated clinical products are necessarily better than general-purpose models.

Imaging

Despite years of “AI radiologist” headlines, essentially every deployed imaging tool is augmentative. The FDA clears most as Computer-Aided Triage, which lets software flag a suspicious scan and move it up the queue (a human still reads it and makes the call).

  • Triage. Aidoc (about $500 million raised, roughly 17 clearances, more than 2,000 hospitals) and Viz.ai, which won the first triage clearance of this kind in 2018, flag strokes, bleeds, and clots; RapidAI and Annalise.ai do similar work across stroke and chest imaging.
  • Reporting. Rad AI drafts the radiology report, which is documentation rather than diagnosis.
  • Pathology. Paige earned the first FDA authorization for AI in the field in 2021 and still only assists the pathologist; it sold to Tempus in 2025 for about $81 million, well under the roughly $241 million it had raised. PathAI is the other main platform.
  • Cardiology. HeartFlow (now public), Cleerly, Eko, and Ultromics quantify plaque or flag low ejection fraction, all as adjuncts a physician confirms.

Even the celebrated firsts here are regulatory-category firsts, not autonomy firsts.

So where does a model make a diagnostic call with no physician in the loop? Across radiology, pathology, and cardiology: essentially nowhere.

The notable exception is diabetic retinopathy screening. Three products, Digital Diagnostics’ LumineticsCore, originally IDx-DR, which in April 2018 became the first autonomous diagnostic AI the FDA cleared, plus Eyenuk’s EyeArt and AEYE Health, diagnose diabetic eye disease from a retinal photo with no eye specialist involved.

The task is narrow and high-volume, specialists are scarce, and there is a billing code (CPT 92229) that pays for the autonomous read.

That last point - dedicating coding - may well be the most decisive factor: Nines, a radiology startup, had FDA-cleared triage and folded anyway for lack of a reimbursement path.

The other place autonomy lives is patient-facing but deliberately non-diagnostic. Hippocratic AI ($3.5 billion valuation, more than 115 million patient interactions) runs autonomous voice calls for post-discharge follow-up and medication adherence. It gets to be fully autonomous precisely because it does not diagnose (which keeps it clear of the device pathway).

Surgery

Surgery: another domain we will gloss over and save for a later piece.

But briefly: today’s RAS is not autonomous. Intuitive’s da Vinci, used in about 2.68 million procedures in 2024, translates a surgeon’s hands at a console into instrument motion. There is always a surgeon with their finger on the trigger.

Real autonomy is still in the lab, if somewhat promising mechanically. A Johns Hopkins robot recently performed part of a gallbladder removal on realistic animal tissue with no human hands, trained on video of surgeons operating. It’s worth stating plainly mechanical precision is not the sole skill that surgeons need - there is deeper judgment and risk calculus at play in an OR. More on this another time.

Diagnostics

Set against the low level of adoption today, the rosiness of diagnostic evidence is a bit curious.

Reported accuracy numbers are high: Autonomous retinopathy screening was cleared at 87 percent sensitivity and 90 percent specificity. An ECG-AI that reads a standard tracing for weak heart pumping validated at an AUROC around 0.93. A pathology foundation model apparently detects cancer across sixteen tumor types at a specimen-level AUC near 0.95.

And, unlike generative AI, some of it has randomized-trial support. In the Swedish MASAI trial, AI-supported mammography found about 20 percent more cancers than standard double reading, at the same false-positive rate, while cutting the screen-reading workload by 44 percent.

In addition to its use in identifying diabetic retinopathy, fundus images can be run through AI models to assess cardiovascular risk factors. They can also identify a patient’s sex from the fundus at an AUC near 0.97 (humans cannot) and help identify early kidney disease.

Beyond eyeballs: There is an FDA-authorized aid for diagnosing autism in toddlers from a questionnaire and a home video. Voice-biomarker tools may be able infer depression from twenty seconds of speech (this evidence is still a bit early and company-reported).

In summary…

The further a task sits from diagnosis, the greater the autonomy afforded (this part is intuitive).

The quadrant that is both autonomous and diagnosis-adjacent remains sparse. Funding follows a similar shape, concentrating where autonomy is safe (e.g., back-office coding), or where augmentation is valuable and unregulated (e.g., documentation and knowledge).

Provider AI by autonomy, proximity to the clinical decision, funding, and adoption A bubble chart. The horizontal axis runs from augmentation on the left to automation on the right. The vertical axis is proximity to the clinical decision, from the back office at the bottom to the diagnosis at the top. Bubble size is the number of startups with more than ten million dollars raised in each segment. Bubble color is clinician adoption, green for high, pale green for medium, gray for low. High-adoption bubbles sit in the augmentation and back-office zones. Imaging and diagnostics is the largest bubble but only medium adoption. The automation-plus-diagnosis corner holds only one small, low-adoption bubble, autonomous retinopathy screening. AUGMENTATION AUTOMATION PROXIMITY TO THE CLINICAL DECISION the diagnosis the back office Imaging & diagnostics (~14) Autonomous dx (~3) retinopathy only Knowledge (~5) Patient-facing agents (~5) Documentation (~6) Back office (~12) Bubble size = startups with >$10M raised ~3 to ~14 Color = clinician adoption high medium low

(Counts approximated based on disclosed funding. Credit to my LLMs.)

Capital and startup volume pile up in imaging and diagnostics - right at the decision - yet adoption remains measured. The high-adoption bubbles (in green) sit off to the side, in documentation, knowledge, and the back office. The one bubble in the automation-and-diagnosis corner is small and gray.

Segment Startups >$10M Adoption
Back office (RCM, coding, prior auth) ~12 High
Documentation (scribes) ~6 High
Knowledge (copilots) ~5 High
Imaging & diagnostics ~14 Medium
Patient-facing agents ~5 Medium
Autonomous diagnosis (retinopathy) ~3 Low

There is clear financial upside here, but the quality of care benefit is more ambiguous.

There is an argument to be made this is too cautious. In many narrow, well-bounded tasks AI seems to match or beat clinicians. At some point the question may not be whether we can trust the machine, but whether we are wise to keep a slower, less consistent human in front of it.

I will play the accelerationist for a moment: there is a real question as to whether slow adoption impedes progress in healthcare AI. Daron Acemoglu and some colleagues published a paper earlier this year about the risk of knowledge collapse from AI - underinvestment in the stock of general knowledge. This is an anthropocentric read - it takes as given that human-legible knowledge assets are the most valuable. In fact, those are exactly the knowledge assets that are the most effortful to consume (reading is hard) and for which it is difficult to scale ingestion.

Creating a corpus of data for RL in medicine has clear and positive future externalities. Training humans is obviously important; training models is as well.

If AI models for patient care are truly close to parity with human physicians, accelerating adoption would be an imperative. And the evidence seems quite strong today - so why the disconnect? A few reasons follow.

Benchmarks =/= evidence

Most AI tools are evaluated on accuracy, time saved, or process compliance. There are some subtle problems with that approach.

The first is that a better process metric is not always a better outcome. In a June 2026 cluster randomized trial at UC San Diego, an LLM scored sepsis-bundle (SEP-1) compliance and gave clinicians near-real-time feedback.

Compliance rose from 70.1 to 82.9 percent (odds ratio 2.10), driven mostly by better documentation of the fluid-bolus step, with 92 percent agreement against human experts. ICU admissions and 30-day mortality did not move.

The second, larger issue is that published accuracy figures are often not representative of real life performance - particularly when moving to a new hospital (this also may explain rational skepticism of diagnostic models). Models often learn shortcuts specific to where they were trained, and many are never tested anywhere else (perhaps federated and swarm learning will save us, but it seems unlikely).

  • A 2018 pneumonia model read the hospital, not just the lungs: it could tell which institution and department produced an X-ray with over 99 percent accuracy, using scanner and formatting cues that happen to correlate with disease, and its accuracy fell when moved to outside sites.
  • Two models that pass identical held-out tests can still behave completely differently once deployed, a failure D’Amour and colleagues named underspecification.
  • The clearest single exhibit is the Epic Sepsis Model. The vendor reported an AUROC of 0.76 to 0.83; on 38,455 hospitalizations, an external validation measured 0.63, with 33 percent sensitivity (missing about two-thirds of sepsis cases), while firing alerts on 18 percent of all patients.
  • Structured reviews routinely judge the great majority of clinical prediction models at high risk of bias with weak external validation, from more than 100 COVID-19 models, none recommended for use, to 62 COVID imaging models, none found clinically usable.

Evidence =/= adoption

You’d expect the market to reward the tools with the strongest outcome evidence, but that doesn’t seem to have been the pattern to date.

There have been 44 randomized trials of AI-assisted colonoscopy showing it finds about a quarter more precancerous polyps (adenoma-detection relative risk around 1.26) than gastroenterologists working alone, and it is still not standard practice, while generative AI with far thinner evidence spread through medicine in about a year. A 2026 review of 4,609 studies of medical language models found only 1,048 used real patient data and just 19 were prospective randomized trials.

Regulation cannot be the only explanation here. Clinicians are given broad latitude to use diverse diagnostic tools as clinical decision aids. The better explanation is that adoption runs on the clinician’s private risk, effort, and benefit calculus, and evidence enters that calculus already discounted - because deployment results so often come in below published ones.

The variables that are practically relevant:

  • Interpretability A note can be verified against the clinician’s own memory of the visit in seconds. A risk score may not be. Clinicians adopt tools whose work they can understand quickly.
  • “Humans-in-the-loop” Scribe output is edited before it commits. A diagnostic nudge, once acted on or dismissed, is hard to reverse.
  • Risk ownership The physician carries the malpractice and moral downside while the vendor and health system capture most of the upside. Under that asymmetry, permission to use a tool is not likely to drive uptake.
  • Economics Scribes have a robust business case, as the billing data above shows. Much of diagnostic AI has no standalone reimbursement path and creates downstream cost for hospitals. AI-assisted colonoscopy, for example: a 2025 microsimulation estimated it prevents about one colorectal cancer per 1,000 people screened but sends about 64 more into surveillance, a cost with no matching revenue, part of why a specialty society declined to recommend it despite the trials.
  • Integration effort A scribe runs in the background and can be switched on by one clinician; a diagnostic model needs imaging-system integration, local validation, and governance sign-off, many more places to say no (in the words of an ML researcher at a major hospital, “it’s hard enough to get the printer fixed”)

The one diagnostic tool that does have strong adoption proves the rule. Autonomous diabetic-retinopathy screening spread because it has its own code, and because autonomous clearance shifts liability away from clinicians.

What could change the wider picture is prospective, multi-site, independent evidence (disclaimer: there is still very little of it). Something like TREWS, a sepsis early-warning system studied across five hospitals and roughly 590,000 patients, where the alerts clinicians actually confirmed were associated with an 18.7 percent relative drop in in-hospital mortality.

What’s coming?

The seed pipeline shows similar bias to the current landscape (if anything, more sharply). Across Y Combinator’s recent batches through Spring 2026, provider-facing healthcare startups cluster heavily in back-office administration: prior authorization, revenue cycle, coding, credentialing, scheduling, referrals, and increasingly voice and phone agents.

In the Winter 2026 batch, all nine provider-facing healthcare companies were administrative automation, with no ambient scribe (mature and saturated) and no diagnostic tools (see above). The names sort into the same few buckets:

  • Prior authorization. Beacon Health (prior auth, referrals, and risk adjustment as “AI employees,” and the batch’s largest healthcare raise), ClaimGlide, and Ruma Care.
  • Billing and revenue cycle. Overdrive Health and LunaBill, both voice agents that call payers.
  • Coding and credentialing. Taiga and Arctic Health, in Spring 2026.
  • Front-office voice agents. Plena Health and Framewise Health.
  • The rare exceptions. Nucleo (oncology CT analysis, and Fall 2025’s one healthcare name on Forbes’ watch list), the Spring 2026 imaging cluster (Adialante’s mobile MRI, Lumius’s 3D ultrasound), and niche scribes (Voquill for pathology, Klarify for therapists).

YC’s own June 2026 Request for Startups names “healthcare administration” explicitly. The framing across the batches moved from “copilot” to “AI agent” and “AI employee,” and the dominant interface became the phone.

Industry observers now describe scribes as functionally commoditized, with more than sixty products and prices that have fallen from about $100 to $60 or $70 a month, and 2026 as the year attention moved to voice agents and claims.

The larger checks here went mostly to slightly earlier or non-incubated companies, such as Latent Health, a 2023 company that raised $80 million to automate prior auths. The 2025 and 2026 batches read as an early, wide bet on healthcare administration rather than a set of breakout rounds.

Who is funding this?

  • VCs (of course). Andreessen Horowitz leads the marquee augmentation rounds (Abridge, Ambience, Hippocratic). General Catalyst pairs investment with owning delivery, having bought a health system, Summa Health, to deploy its portfolio’s tools inside a real provider. GV made OpenEvidence its signature position, and Nvidia’s venture arm recurs as a strategic across Abridge and Aidoc.
  • Health systems. Abridge was built inside UPMC. Mass General Brigham spun out CodaMetrix. Memorial Sloan Kettering spun out Paige. Mayo is co-developing models in cardiology and beyond.
  • Epic. The incumbent across hospitals with ~300 million patients, it is the quiet center of gravity. Also, increasingly a competitor to startups that build on top of it.

Where it goes

The safe (and valuable!) middle, scribes and copilots, is growing saturated.

For the rest: the most determinative open questions are less about model architecture than about the required financial and technological infrastructure. A billing code that reimburses autonomous actions, a regulatory path for continuously-learning / non-static models, training and data paradigms that improve external validation.

More to come.


← All writing