2026 · 09 · 25 · 8 MIN READ · FIELD NOTES · ISAAC USIFO, ROSELINE CHIMA-KALU

Building a more reliable health assistant for Nigeria

What early testing taught us about local knowledge, clinical boundaries, and earning trust.

We have completed the early testing phase for Alma Assistant V2, our conversational health assistant being developed for the Nigerian context. V2 focused on grounding responses in selected clinical guidance, improving local medicine lookup, and strengthening controls around how answers are produced. Those findings now inform the next version, which is under development.

Our ambition is to create a connected experience in which people can understand their health, manage it over time, and work more effectively with those providing their care.

Alma Assistant, our conversational health product, is one part of that broader effort. This article shares what we learned while developing its second version for the Nigerian context.

Our goal is to build a platform people can rely on and healthcare professionals can trust. That sets a demanding standard for the assistant. An answer needs relevant evidence, appropriate boundaries, and a clear account of what the system can and cannot establish.

The first version established the conversational experience: people could ask health questions through a web interface and receive responses from underlying language models. V2 focused on the systems behind those responses, including retrieval from selected clinical documents, medicine lookup, and stronger controls around how answers are produced.

We recruited 14 doctors for early testing, of whom 10 actively participated. Alongside the feedback received, we reviewed conversations, investigated failures, and ran targeted engineering checks.

This was early product testing, not a clinical validation study. Its value was in showing us where the product’s behavior fell short of its intentions and what needed to change.

Several lessons now guide how we build Alma Assistant.

Local knowledge has to reach the answer

A health question often carries context that is easy to miss: a medicine’s local brand name, the way someone describes a symptom, or the guidance used in their healthcare setting.

V2 introduced a retrieval layer that searches selected clinical documents and supplies relevant passages to a language model before it generates an answer. These include Nigerian national guidance and international reference material. A separate lookup uses NAFDAC product-register data to help identify medicines by brand or generic name.

The underlying models help interpret questions and generate responses. Alma’s surrounding systems select information, assemble context, and apply additional controls to parts of the answer.

Building those connections exposed an important distinction: possessing the information is not the same as using it correctly.

Someone may type a short brand name, while the registry stores a longer product title that includes its strength and dosage form. A lookup can miss that product entirely. It can also fail in the opposite direction, matching an ordinary word in a question to an unrelated product.

We improved bare-brand recognition, tightened matching around word boundaries, excluded medical-device and veterinary categories from medicine results, and made result ordering consistent. We also corrected a data-loading failure that could temporarily disable much of the drug-name detection.

These changes improve a specific part of the system. They do not make medicine identification a solved problem. Formulations, combination products, and ambiguous names still require better handling. Our next work includes updating the underlying lexicon and evaluating those difficult cases more systematically.

Registry data also has limits. Registration does not establish that a medicine is currently stocked nearby, suitable for an individual, or available at a particular price. Those questions require different evidence.

Clinical boundaries need more than instructions

Some of the most consequential findings concerned how far an answer should go.

Reviewed conversations included responses that became too specific about treatment plans and dosing without adequately establishing the individual context. They also showed why a user should not have to challenge an answer before the system recognizes an important limitation.

We added controls for unsolicited dosing and for users who identify themselves as minors. We also introduced a boundary around composing patient-specific treatment regimens.

Testing that boundary produced a useful engineering lesson. A detector could correctly identify a request while the underlying model still generated the kind of answer the instruction was intended to prevent. Different model providers did not follow the same instruction equally reliably.

We therefore added a check before certain responses are shown. For requests identified as crossing the prescribing boundary, Alma checks the generated answer for specified regimen patterns. When it detects a violation, it replaces the response with an explanation of the boundary and the general information it can provide.

A subsequent spot check also returned a medication schedule in response to a treatment-plan request. This shows that the intended boundary is not consistently enforced in the deployed experience; identifying the failing path and testing its coverage remain necessary.

The protection remains dependent on what the checks recognize. It is not a guarantee against every unsafe response, and it does not establish comprehensive paediatric safety. These checks address a defined failure mode, but their coverage and consistent enforcement still require evaluation.

A source needs to support the trust placed in it

People reasonably expect a source shown beneath a health answer to support what they have just read.

Our work on citations showed how easily a system can fall short of that expectation. A document may resemble the question without supporting the answer. A useful passage may reach the model but fall below the threshold for displaying a source. A short clarification can inherit the clinical context of an earlier message and receive a citation it does not need.

We introduced a citation filter, corrected response paths that bypassed its result, and added instrumentation to record which sources were emitted with completed answers. We also revised the general-knowledge badge so greetings and clarifying replies are less likely to receive misleading labels.

These changes make the behavior more deliberate and more observable. They do not yet establish support for every individual claim. The current citation check uses textual overlap, which cannot fully establish whether a source supports the clinical meaning of an answer.

That remains an important part of our next evaluation work. We need to assess whether the evidence supports the recommendation, not simply whether retrieval found a related document.

Reliability includes finding care and completing a conversation

The work also extended beyond clinical answers.

A person searching for a named diagnostic centre needs the search to consider its name. Our earlier facility-search approach relied too heavily on proximity. We changed the search method, tightened location handling, and corrected map behavior that could produce results unrelated to the facilities returned.

We improved photo uploads, made screenshot-upload failures visible in the feedback flow, and strengthened continuity when an underlying model service was unavailable.

We have increasingly tested these paths by their observable results, including whether our own checks can detect a known failure.

What the work has not established

The V2 testing phase produced targeted engineering evaluations and recorded examples of improvements. It did not establish overall clinical accuracy through a standardized, clinician-reviewed benchmark. Completing that phase is a development milestone, and broader claims about reliability still require further evidence.

Our knowledge coverage is also uneven. A corpus audit distinguished material available to clinical questions, surveillance material, documents that current retrieval paths cannot reach, and reports excluded because their dates were not parsed reliably. That accounting gives us a more useful expansion plan than a single headline document count.

The local symptom-idiom lookup is built, but its entries remain in draft pending review.

Privacy is another area where implementation and public commitments must agree. Our review identified remaining work in sensitive-query logging and retention handling.

These findings define work still to be completed. They also help us make more precise claims about what the product can do today.

The next stage

Work on V3 is now underway, informed by these findings. Our focus is to expand the knowledge base while making reliability measurable across the whole answer: interpreting the question, retrieving appropriate evidence, using it correctly, respecting clinical boundaries, and communicating uncertainty.

We are developing stronger methods for selecting relevant evidence and checking how it supports an answer. Alongside that work, we plan to update the medicine lexicon, improve coverage, and build a clinician-reviewed evaluation set that can be run repeatedly as the system changes.

Progress will need to be demonstrated across different questions, conversation patterns, and model providers. Privacy and operational verification belong in that evaluation too.

We want future performance claims to be accompanied by their scope, methods, and limitations. We also want the feedback loop to remain open: the answers that need correction are among the most useful inputs to the product.

The doctors who participated in early testing helped us examine Alma Assistant against real expectations of a health-information tool. We are grateful for their time, feedback, and the questions they raised.

The assistant is one part of Alma’s broader work toward intelligent healthcare for Africa. What we learn here informs how we develop and evaluate other tools across consumer health management and provider workflows.

Our ambition is to earn a place as a trusted platform for both consumers and healthcare professionals. V2 is a step toward that ambition. The work ahead is to keep demonstrating, through evidence, where that trust is earned.

Alma Assistant supports access to health information. It does not replace an individual assessment by a qualified healthcare professional.

The V2 early-testing phase has concluded, and new registrations are currently closed as development continues. If you are a clinician, researcher, or institution interested in contributing to future evaluation or Nigerian health-information coverage, we would welcome a conversation.