Care that answers first.
Free health guidance in any language, at any hour. No appointment, no account, no fee. When a person needs a clinician, this reaches one.
First Voice Health is building the evidence base for AI triage in the languages where it is needed most.
Key statistics
- 4.5 billion people lack coverage for essential health services (WHO and World Bank, Universal Health Coverage Global Monitoring Report 2023)
- 1M+ healthcare providers use doxy.me, the world's largest free telemedicine platform. (internal platform data, 2026)
- 51.6% of true emergencies undertriaged by a leading consumer health AI, in English, in a structured clinical test (Ramaswamy et al., Nature Medicine, 2026)
Nobody has measured what happens when the AI is the only thing there.
Every published evaluation of medical AI measures it helping a clinician who is already in the room. A randomized study of 9,691 patients across 16 Kenyan primary care clinics measured AI assisting a clinician. A five million dollar research program has benchmarked AI support for frontline health workers across Kenya, Nigeria, and Rwanda. All of it assumes a trained human is present.
That assumption does not hold for most of the world. It does not hold for a mother in Kano at two in the morning, forty kilometers from a clinic, with a feverish child and a phone.
A structured clinical test of a leading consumer health AI found it undertriaged 51.6 percent of true emergencies. The same test applied to a tool built for physicians found 12.5 percent (Ramaswamy et al., Nature Medicine 2026, and a replication on a physician-facing system, medRxiv 2026). In a physician-rated study across six languages, the clinical substance of chatbot health answers varied nine times more by the language of the question than by which chatbot answered it (Ariel et al., medRxiv 2026).
One study comes close to ours. In June 2026, a symptom checker inside South Africa's government maternal health platform was rated safe in 98.4 percent of cases by an independent physician panel (Ada Health, June 2026). It ran in English, it covered maternal health only, and the company that built the tool sponsored the evaluation.
What is missing: an independent measurement of a language model answering a patient who has no clinician available, for general health problems, in a language the safety research has never covered. That is the condition we intend to serve, so measuring it is our obligation before it is our opportunity.
This is already happening.
A voice AI service answered roughly 614,000 questions from 36,000 people in Zambia on basic phones with no internet, and nearly a quarter of those questions were about health (GSMA and Viamo, Zambia case study, 2025). It has since launched in five more countries. No clinical safety evaluation of it has ever been published.
We are not proposing to build something risky and measure it first. The risky thing shipped three years ago and is scaling. We are proposing to measure what is already reaching people.
Sources: 51.6 percent and 12.5 percent, Ramaswamy et al., Nature Medicine 2026, and a replication on a physician-facing system, medRxiv 2026. Six-language study, Ariel et al., medRxiv 2026. Five million dollar program, Gates Foundation funded benchmarking of AI assistance for frontline health workers in Kenya, Nigeria, and Rwanda. 98.4 percent, physician panel evaluation of a maternal health symptom checker in South Africa, Ada Health, June 2026. 614,000 questions, 36,000 users, share of health questions, and five-country expansion, GSMA and Viamo, Zambia case study, 2025. 9,691 patients, randomized study across 16 Kenyan primary care clinics.
The number to beat is 3.7 percent.
A triage system that is 90 percent accurate can still be lethal, because the failures are not evenly costly. Sending someone to a clinic they did not need wastes money. Telling a mother her child's fever is nothing, when it is malaria, kills the child. So the measure is undertriage, not accuracy.
The honest comparison is a human doing the same job by phone. In a study of 1,294 recorded out-of-hours calls, nurses using computerized decision support undertriaged 3.7 percent of cases as judged by an expert panel. Physicians working the same calls without the tool undertriaged 7.3 percent. Our target is to beat the nurses. We will publish the number whether or not it flatters us. (Graversen et al., BMC Family Practice, 2020)
The First Work
- Collect 800 real de-identified health questions from people already using a phone-based service, under formal data-sharing agreements.
- Have each question adjudicated independently by three in-country clinicians for what it actually needed: emergency, same-day, routine, or self-care. A fourth clinician resolves disagreements.
- Test every available model against the set in four arms: a frontier model and a small model that can run offline, each in English and in the target language. Report undertriage by language, by age band, and by presentation.
- Publish the questions, the harness, the protocol, and the results openly, including to competitors.
The first language is Hausa.
Why Open
The benchmark is a public good or it is worthless. A safety standard that one organization owns is a marketing asset. Every implementer, ministry, and model developer should be able to run this and see their own number. We are funding the work. We are not keeping it.
Who Is Behind It
First Voice Health was initiated by Dylan Turner, co-founder of doxy.me. doxy.me is used by more than one million healthcare providers and has carried over 490 million patient visits (internal platform data, 2026).
The safety layer is not the model.
Deterministic rules run before the model and again after it. A rule can raise the urgency of an answer. A rule can never lower it. Every rule that fires is visible, with the text that triggered it and the rule that caught it.
When a family member minimized symptoms in the Nature Medicine test, the AI shifted its advice toward less urgent care. A rule that can only escalate is what stops that.
Who this is with.
We are looking for programs that already reach people by phone, in their own language, at scale. If you run one, the questions your users already ask are the benchmark. Partner with us at /partners-inquiry.
Adjudicate the first benchmark.
We are recruiting practicing clinicians to independently adjudicate real de-identified patient questions and decide what each one needed: emergency, same-day, routine, or self-care. Adjudicators are paid and are named co-authors on the published result. This is not a request for funding and not a commercial partnership. Sign up at /adjudicators.