MsingiAI
Back to News
Swahili Voice AI, an Ordinary Phone Call Away
Announcements

Swahili Voice AI, an Ordinary Phone Call Away

Published September 14, 2026 · MsingiAI Editorial Team

During a week of field testing in Busia and Ahero, someone asked whether a human was speaking on the other end of the line.

They were talking to Nia, a voice AI agent.

For our co-founder and Voice AI Research Engineer, Kiplangat Korir, the response was a welcome surprise. He had gone into the field expecting disappointment, conscious of the many ways speech systems can struggle outside controlled conditions. The conversations that followed left him more optimistic about what the technology could make possible.

Working alongside his colleague Betty Kyallo and the wider team, Kiplangat joined farmers from Hello Tractor to deploy and test the Nia Voice AI Agent under the AI Hub for Sustainable Development, with technical guidance from Crane AI Labs.

At MsingiAI, the experience connects closely to our work on Sauti, our Swahili speech recognition and speech generation models. It gave us practical lessons about how people interact with voice AI, what they expect from it, and where our engineering effort needs to go next.

The starting point was remarkably familiar: a phone call.

Farmers could reach the agent by making an ordinary voice call using airtime. They could speak in Swahili and switch naturally into English. On the caller’s side, accessing the service did not require installing an app or buying a mobile data bundle.

That is a concrete step towards democratizing AI. It brings a new capability into something people already know how to do. The first instruction can be as simple as calling a number.

There is significant engineering behind that interaction. Speech recognition, language processing, speech generation, and telephony all have to work together. For the person making the call, the experience should be straightforward: ask a question, be understood, and receive a useful response.

The fieldwork made the importance of that simplicity clearer. It also showed how much people expect once they can reach the technology.

Farmers discussed reliability, access, and equality. These are substantial expectations, and they give us a useful standard for development. Someone trying an AI agent is entitled to care about whether it works consistently, whether they can afford to use it, and whether it understands them as well as it understands someone else.

Reliability extends across repeated conversations. A good first call needs to be followed by a good second call. The service needs to remain useful when someone changes their phrasing, asks a follow-up question, or needs to correct a misunderstanding. We want our evaluation to reflect that experience over time.

Access also deserves careful attention beyond the ability to dial a number. Airtime costs money. Long pauses and repeated clarification can add to the cost of a conversation. Clear responses, sensible call duration, and reliable availability all affect whether the service remains practical to use.

Equality becomes tangible when we examine whose speech a system understands. Different accents, local expressions, and levels of confidence with technology can produce different experiences. We need to look closely at who completes an interaction comfortably and who has to keep repeating themselves. That detail matters alongside overall model performance.

For MsingiAI, these lessons reinforce the purpose of our work on Sauti. We are building speech technology around the languages and speaking patterns people bring to everyday life.

Kenyan conversations move naturally between Swahili and English. People use familiar expressions and shorthand. They change direction halfway through a sentence. They do not pause to organize their thoughts around a model’s vocabulary.

The technology has to get better at following them.

That shapes the data we work with, the examples we evaluate, and the errors we prioritize. An almost-correct transcript can still miss the detail that determines the right answer. We need to understand what recognition errors mean for the whole conversation and make it easier for the agent to clarify something it has not understood.

There were also practical challenges in the field, including some background noise and occasional hallucinations. These provide specific areas for improvement.

For noise, the next steps include better audio handling, testing under representative conditions, and clearer recovery when speech is difficult to hear. For hallucinations, we need to examine how unsupported answers arise, improve the use of dependable information, and strengthen clarification and uncertainty handling. These are addressable areas of work, with improvements to be checked against the cases that exposed them.

The question about whether a human was on the line makes that last point especially relevant. A natural voice can make an interaction feel comfortable. We need the quality of the answer to justify the confidence its delivery can inspire.

We also want to measure more of what a caller experiences. Did they get useful help? Was the response understandable? Could they correct an error without starting over? How much time and airtime did it take?

These questions help connect model development to practical outcomes. Field observations give us test cases. Recurring difficulties help us set priorities. Feedback from farmers helps us see which improvements will make a difference to the person using the system.

As work continues towards UNGA 2026, we are encouraged by what happened in Busia and Ahero. The response exceeded Kiplangat’s expectations and gave us a clearer picture of both the opportunity and the work ahead.

Our ambition at MsingiAI is to make African language speech technology useful in everyday life. An ordinary phone call is a powerful place to begin. People can approach the technology with a question, a familiar language, and a device they already use.

We want to keep building towards a future in which that call reliably connects them to useful help.


We thank the farmers who shared their time and feedback, Hello Tractor, Betty Kyallo, Mildred Rebecca Namagembe, Kato Steven Mubiru and Crane AI Labs, and everyone involved in the fieldwork. We are grateful to the AI Hub for Sustainable Development, Keyzom Ngodup Massally, Raiyan Arshad, Dwayne Carruthers, EkStep Foundation, CINECA, and Africa Compute Fund for supporting this work.