2026 · 08 · 23 · ANNOUNCEMENTS · KINDRED

Alma Web V2: answers grounded in Nigerian clinical guidelines

A general model knows medicine in the abstract. Alma V2 retrieves what Nigeria's own health authorities actually published, cites it, and refuses when it cannot find it.

Alma web V2 is in testing. It answers health questions for people in Nigeria and West Africa, and the thing that makes it different from a chat window over a general model is what happens before it answers: it goes and finds what Nigeria’s own health authorities published, and it tells you what it found.

The gap is not intelligence. It is location.

Most Nigerians get health information from a search box, a WhatsApp group, or a conversation across a chemist’s counter. Frontier models are extraordinary at medicine in the abstract — and that abstraction is exactly the problem. Ask one about hypertension management and it will answer well, for a patient in a country whose guidelines it absorbed in bulk. It will not know which drugs are actually registered and stocked here, which brand name maps to which generic in a Lagos pharmacy, what the national protocol says when it diverges from the international one, or what somebody means when they describe a symptom in the words people actually use.

This is not a hallucination problem in the usual sense. The answers are frequently correct in a vacuum and wrong in a clinic in Ibadan. On a health surface that distinction is the entire product.

We built V1, watched it, and found four failure modes worth naming. It reached for generalised training knowledge over local guidance. It could not tell a user whether a claim came from a verified document or from the model’s own reasoning. It handled sensitive questions — particularly around sexual and reproductive health — without the care that category demands. And it was brittle: when its primary model provider failed, the whole thing failed with it.

V2 is the answer to those four.

Alma open on a phone, greeting the reader by name under the question “How can I help with your health today?”, above prompts for checking symptoms, understanding medication, explaining lab results, pregnancy and maternal care, finding care nearby and a daily health tip. A footer reads that Alma provides health information, not medical advice.
Alma web, V2

Retrieval first, generation second

Alma V2 does not ask a model what it knows. It searches a curated corpus of Nigerian health documents, assembles what it finds, and only then asks a model to explain it.

Three sources feed that context. The first is the clinical corpus — national treatment guidelines, public health surveillance, and the reference literature underneath them. The second is a drug register built on NAFDAC data, which resolves the brand name someone actually has in their hand to a generic, a dosage form, a strength, a manufacturer. The third is a symptom lexicon intended to map local idioms — the words people use for illness — onto clinical terms.

When a follow-up question is short enough to be ambiguous on its own — someone simply typing “and in pregnancy?” — the system searches both the bare question and the question carried forward with its conversational context, then keeps whichever retrieval came back stronger. Short questions are how people actually talk, and losing the thread on them was a real defect.

Citation, and the honesty of a miss

Every claim Alma makes above a similarity floor carries its source, rendered under the answer. That much is table stakes.

The part we care more about is what happens on a miss. When a clinical question finds nothing in the corpus, Alma does not quietly fall back to what the model absorbed in training and present it in the same voice. It marks the answer as general medical knowledge rather than verified Nigerian guidance, and says so on the face of the message.

And at the end of the provider chain, if a question is clinical and the retrieval came back empty, the system is configured to fail closed — to return a refusal rather than a confident guess. A wrong answer about a child’s fever is worse than no answer. That sentence is a design constraint, not a slogan.

Where the retrieved documents disagree — where a national guideline and an international one recommend different things — Alma presents both side by side rather than silently picking a winner. Which guidance a clinician follows is a clinical judgement, and it is not ours to make invisibly.

What is actually in the corpus

Alma V2 searches 12,683 actively retrievable document chunks, drawn from named authorities: the Federal Ministry of Health of Nigeria, the Nigeria Centre for Disease Control, the World Health Organization and WHO AFRO, the DHS Program with NPC Nigeria, and the Diabetes Association of Nigeria with the Nigerian Society of Endocrinology and Metabolism. The drug register holds 8,791 entries — 989 brands mapped across 102 generics.

A further 3,102 chunks of surveillance data sit outside the searchable set because their reporting weeks could not be parsed reliably. Undated epidemiological data is not safe to reason from, so it is excluded until it can be dated. We would rather have a recency gap we can describe than a confident answer about an outbreak drawn from a document we cannot place in time.

Privacy, where it costs something

Questions about sexual and reproductive health are redacted from our own retrieval logs before they are written. We keep audit records of what the system retrieved so we can investigate failures, and this is the category where that instinct has to yield: the log records that a retrieval happened, not what was asked. The model is also instructed to answer those questions with more care and no judgement.

Medical records are append-only. A correction is a new record that points at the one it amends; nothing is overwritten, and nothing is deleted. A record that can be quietly edited is not a record.

What we have not done

Alma V2 is in testing, and the honest list of what is missing is short and load-bearing.

No clinician verification at scale. Alma’s outputs have not been systematically reviewed by licensed medical professionals. Nigerian clinicians are testing it now; that work is not finished, and we will not describe it as finished.

No benchmark. There is no standardised evaluation measuring accuracy, hallucination rate or citation fidelity against a fixed set of clinical queries. Until there is, any number we quoted about how good Alma is would be a number we made up.

The symptom lexicon is not live. Every entry in it is still marked draft, and the module only accepts reviewed or verified mappings. So local idioms are not yet being translated to clinical terms in production. The mechanism is built; the reviewed content behind it is not.

Coverage gaps. Less common conditions and specialised protocols are not yet digitised into the corpus.

What comes next

A clinical benchmark, and then that benchmark as a gate on every model and prompt change we ship afterwards. A clinician review panel working through the hard cases — the conflicting-guideline ones especially. Auditing the symptom lexicon up from draft so the idiom layer can switch on. Expanding the corpus into the gaps, and recovering the surveillance data we had to set aside.

We will publish an evaluation before we publish a claim about accuracy. Until then, what we can say is what Alma is built to do, what it refuses to do, and what it does not yet know.

Alma is at almahealth.xyz. If you are a clinician, a researcher, or an institution with guidelines that should be in the corpus, we would like to hear from you.