When there is no doctor, people ask anyway.
They are asking AI right now, in languages nobody has tested, and nobody is checking whether the answers are safe. We are measuring that first. Then we are building the thing that answers, and reaches a clinician when the answer is not enough.
The problem
4.5 billion people are not fully covered by essential health services (WHO and World Bank, Universal Health Coverage Global Monitoring Report 2023).
The number hides the ordinary version of itself. A mother at two in the morning with a feverish child. The clinic is forty minutes away and costs a day of wages. She has a phone in her hand and nobody to ask.
So she asks the phone. That already happened, quietly, over about three years, and nobody planned it.
Why this is bigger than it looks
The answers are already being given, and they fail in one direction.
Ramaswamy and colleagues at Mount Sinai published in Nature Medicine in February 2026. Across 60 emergency cases and 960 model responses, ChatGPT's health mode undertriaged 51.6 percent of the emergencies. When the question carried a cue toward a milder reading, the advice followed it, odds ratio 11.7. A preprint in April ran the same test against a tool built for doctors and got 12.5 percent.
Nobody has run that measurement on a service people actually use, in the language they actually speak, on questions they actually asked.
A correct answer is worth less than it sounds.
Ada Health's decision support ran inside South Africa's MomConnect. An independent panel of three senior physicians judged 184 cases and rated the advice safe 98.4 percent of the time. The same paper reports that the advice matched the panel's chosen urgency in 48.9 percent of cases and was more cautious in 45.1 percent. Safe counted caution as success.
Then the part that matters. After receiving advice rated safe, user undertriage stayed at 51 percent of decisions, and one in four women still acted less urgently than she had been told. Lower-income participants were likelier to act unsafely despite safe intentions, because something outside the conversation stopped them. Money and distance were the binding constraint, not advice quality.
A service that works can still disappear.
Babyl registered 2.8 million Rwandans and integrated with 450 of Rwanda's 510 primary health facilities. Its UK parent went into receivership, its US subsidiary filed Chapter 7 on 9 August 2023, and the Rwandan service wound down that September. A study in BMC Primary Care found facility visits afterward ran 15 to 22 percent above the pre-Babyl baseline, with 809 more respiratory infection cases and 256 more malaria cases a month. That is an interrupted time series, so read it as demand with nowhere to go rather than proof of clinical effect. The lesson is about money. The clinical layer was not what failed.
Sources: Ramaswamy et al., Nature Medicine, 23 February 2026. OpenEvidence comparison, medRxiv 2026.04.23.26351526, preprint. Ada SAFEMOM, Nature Health, 16 June 2026. Babyl, BMC Primary Care, January 2026.
What we are measuring, before we build anything
Patient-facing health AI is being evaluated in fragments. One language pair here, one condition set there, mostly by the people who built the system, almost always on invented cases.
A randomised study across 15 Kenyan clinics measured AI helping a clinician who was already in the room. A Gates-funded programme has benchmarked AI support for frontline health workers across three countries. Both assume a trained human is present.
Two studies in 2026 came closer. A symptom checker inside South Africa's government maternal health platform was rated safe in 98.4 percent of cases by an independent physician panel, in English, for maternal health only. A WhatsApp triage service in South Africa was tested across isiZulu, isiXhosa and Afrikaans and reported 3.3 percent undertriage with no missed emergencies, using invented cases and a gold standard set by its own developer.
Both are real work and we cite them. Neither answers the question we are asking. What happens to a person with no clinician to reach, asking in their own language, when the only thing that answers is a language model.
Nobody has run a pre-registered, independently judged evaluation of a deployed service's real patient traffic, reporting both kinds of error. That is the study.
Sources: 15 Kenyan clinics, randomised study of AI assistance in primary care, 2026. Gates Foundation funded benchmarking of AI assistance for frontline health workers in Kenya, Nigeria, and Rwanda. 98.4 percent, independent physician panel evaluation of a maternal health symptom checker in South Africa's government platform, 2026. 3.3 percent, developer-run evaluation of a South African WhatsApp triage service across isiZulu, isiXhosa and Afrikaans, 2026.
Two kinds of error, and both get reported
A triage system that is 90 percent accurate can still kill people, because the failures are not evenly costly. Telling a mother her child's fever is nothing, when it is malaria, can kill the child. That is undertriage and it is the error that matters most.
A system judged only on undertriage can score perfectly by sending everyone to hospital. That is overtriage, and for a family who cannot afford the trip it is its own harm. A safety number that counts overcaution as safe is not measuring safety.
So we report both, side by side, every time. We are not publishing a target we promise to beat. We publish what we find, next to what human telephone triage achieves, and let the comparison stand. Published undertriage in human telephone triage ranges from about 2 percent to 19 percent depending on setting and definition.
How the answer reaches a person
A phone number, not an app. WhatsApp, SMS, USSD and voice, plus a web page small enough to load on a 2G connection. The people this is for are on feature phones. A service that needs a smartphone and a data plan has already excluded them. Four doors, one clinical brain, the same safety rules on all of them: Web, SMS, Voice, and USSD.
The urgency decision does not belong to the model
Deterministic rules run before the model and again after it. A rule can raise the urgency of an answer. A rule can never lower it. Every rule that fires is shown, with the text that triggered it and the rule that caught it.
This pattern is not ours. A South African team published a version of it in April 2026, 53 clinical rules coded across 11 languages, running independently of the model. We cite their work and we build on it.
In the Nature Medicine test, a cue toward a milder reading pushed the AI's advice toward less urgent care. A rule that can only escalate is what stops that.
53 rules and 11 languages, safety architecture of a South African WhatsApp triage service, published April 2026. Cue effect, Ramaswamy et al., Nature Medicine, 23 February 2026.
Who pays for this, and why that is a problem we take seriously
First Voice Health was started by a co-founder of doxy.me, and doxy.me pays for it. That is a conflict of interest. Naming it is more useful than denying it.
So the rules are fixed in advance. The clinicians who judge the results have no commercial relationship with doxy.me or with any partner supplying the data. The protocol and the analysis plan are registered publicly before we see a single result. An independent statistician who does not work for us runs the numbers. The study design is frozen and published before we build any product in the category we are measuring.
We publish the result whether or not it flatters us, including a result that ends this organisation's reason to exist.
Answering is the easy half. The reason this is worth building is that the second half is reachable. doxy.me removed the download, the login, the account, the institution and the invoice from a video visit, and a patient joins by tapping a link. In the Philippines, 527 clinicians on that platform took calls in the last 90 days. That is what makes "you need to see someone" into something other than a shrug. (doxy.me platform figures, internal, 11 August 2026.)
The first study
- Take 400 real health questions people have already asked a live phone-based service, de-identified under a formal agreement. Real questions, not invented ones.
- Have each question judged independently by three in-country clinicians for what it actually needed: emergency, same-day, routine, or self-care. A fourth resolves disagreements. None of them work for us or for the partner.
- Ask the same 400 questions of the deployed service, of a frontier model, and of a small model that can run offline, each in English and in the local language. Report undertriage and overtriage for every one.
- Have local clinicians answer the same 400 questions by text, so the AI has a human comparison from the same setting rather than from a different continent.
- Publish the questions, the method, the analysis plan and the results openly, including to competitors, and register all of it before the first result exists.
The language and the country follow the first partner. We measure where people have no clinician to reach.
Why Open
The benchmark is a public good or it is worthless. A safety standard that one organisation owns is a marketing asset. Every implementer, ministry, and model developer should be able to run this and see their own number. We are funding the work. We are not keeping it.
Who this is with
We are looking for services that already answer health questions at scale, in people's own languages, on the phones they already own. If you run one, the questions your users are already asking are the study. We evaluate your service independently, publish the result, and name your team as co-authors. You will not have a veto over what we find. Partner with us at /partners-inquiry.
Adjudicate the first study
We are recruiting practising clinicians to judge real de-identified patient questions and decide what each one needed: emergency, same-day, routine, or self-care. Adjudicators are paid and are named co-authors on the published result. To keep the study independent, adjudicators cannot have a commercial relationship with doxy.me or with the partner supplying the questions. This is not a request for funding and not a sales conversation. Sign up at /adjudicators.