Study finds 98.8% of Alodokter AI consultations clinically safe, independent review shows

A new study led by Stanford researchers and co-authored by UC Berkeley professor Ziad Obermeyer has found that 98.8% of a sample of Alodokter AI’s real-world telemedicine consultations were rated clinically safe following independent physician review.

Alodokter AI is developed by Alodokter, an Indonesian digital health company founded in Jakarta in 2014 by Nathanael Faibis and Suci Arumsari, with 20 million monthly active users. Deployed since January 2026, the AI operates as a background support tool for doctors during chat-based teleconsultations, assisting with history-taking, diagnosis, medication and testing recommendations, triage and referral. A licensed physician reviews, can amend or reject its suggestions, and remains the sole prescriber.

Three independent Indonesian physicians, with no prior employment, freelance or financial ties to Alodokter, each assessed 402 de-identified consultations completed between 19 and 21 April 2026 against 14 safety criteria spanning information-gathering, red-flag detection, diagnosis, medication and dosing, testing, triage, referral, safety-netting and unsupported AI content.

Under the study’s consensus rule, 397 of the 402 were classified clinically safe: 388 as safe and nine as safe with a minor, non-critical omission. Five were classified unsafe. The sample deliberately over-represented emergency, referral and other higher-acuity cases, and the 95% confidence interval for the result was 97.1–99.5%.

Researchers say the study is distinctive for combining real-world deployment, end-to-end consultation assessment and independent per-case physician review — features rarely evaluated together in clinical AI research to date. Earlier company-authored studies from telemedicine AI providers Curai and Doctronic have measured agreement with clinicians on diagnoses or treatment plans, but did not assess the full consultation pathway or involve independent review of every case.

Obermeyer said “real-world consultations are complex and unpredictable.”

Alodokter’s chief medical officer, Dr Louise Hewitt, said turning large language model capabilities into reliable clinical support “requires deep clinical experience and rigorous evaluation.”

Stanford doctoral researchers Wendy Yin and Natalia Khoudian, who worked on the study’s design and methodology, said clinical AI should be judged in the healthcare systems where it will actually be used.

Alodokter’s chief executive, Nathanael Faibis, said the tool draws on the company’s decade of experience running large volumes of teleconsultations.

According to Alodokter, adoption of the AI tool among its general practitioners is now approaching 100%, and the company reports that 98% of GP-led consultations supported by the AI received positive patient satisfaction ratings.

The study’s authors note several limitations: it was a retrospective, single-platform assessment using a deliberately higher-acuity sample, reviewer judgements varied, and no control group was used to compare patient outcomes.

Alodokter says the five cases classified unsafe point to concrete priorities for further work, including escalation of acute warning signs and paediatric dosing checks. The company is now working with the same research team on a randomised controlled trial to assess the AI’s impact on clinical quality, doctor productivity and patient outcomes.

The study, Assessing the Clinical Safety of an AI Assistant in Indonesian Chat-Based Telemedicine, is published as Stanford King Center on Global Development Working Paper No. 2078 (19 August 2026).

Author


Discover more from HealthTechAsia

Subscribe to get the latest posts sent to your email.