MsingiAI

Publications

Peer-reviewed papers, technical reports, and benchmark studies from MsingiAI Labs.

6 Papers
Topics:
Technical ReportTechnical ReportJune 2026

AkiliCode-14B: Repair-Oriented Post-Training for an Open 14B Code Model Technical Report

MsingiAI

We present AkiliCode-14B, a research-preview code model obtained by post-training Qwen2.5-Coder-14B through a three-stage pipeline: (1) instruction tuning, (2) execution-aware reinforcement learning, and (3) failure-focused repair training. The final stage is explicitly designed to teach code revision from failed attempts rather than to solve random easy problems from scratch. Relative to our best Stage 2 checkpoint, the promoted Stage 3 checkpoint improves HumanEval+ from 60.98 to 62.80 and MBPP+ from 65.08 to 65.61, while BigCodeBench-Instruct decreases from 45.61 to 45.09. On CRUXEval-O, the promoted checkpoint reaches 49.75. On a Live-CodeBench v6 official-style reimplementation over 1,055 problems, the same checkpoint obtains 11.37 single-greedy pass@1, with 100.0% extraction success and 98.58% syntax-valid outputs. The dominant failure mode on LiveCodeBench is wrong answers rather than parsing, syntax, or extraction failure. Our main conclusion is mixed but useful: repair-oriented post-training improves local function-level synthesis in this run, but does not close the gap on hidden-test algorithmic generalization. We release AkiliCode-14B as a research preview and use this report to document both what worked and what did not.

Technical ReportTechnical ReportMay 2026

Sauti ASR Technical Report

MsingiAI

We present Sauti ASR, a dual-track automatic speech recognition (ASR) effort for Swahili designed to separate release readiness from research exploration. Track A is a pragmatic release lane built on top of Microsoft's Kenyan-focused Paza Whisper checkpoint, while Track B is a research-preview lane based on Meta FAIR's Omnilingual LLM-ASR stack. Track A is the current official public release and achieves 13.72% word error rate (WER) and 3.88% character error rate (CER) on a 500-sample held-out Kenyan Swahili benchmark. Track B is not yet the production-default model, but its tuned step-250 checkpoint reaches 15.13% WER on a 1,000-sample development slice, improving over an earlier 16.36% internal baseline. The value of this report is not a state-of-the-art claim. Instead, it is a technically candid account of model selection, tuning dynamics, inference-path divergence, long-form failure modes, and public release constraints. We show that benchmark quality and long-form usability are not the same optimization target, that best-checkpoint selection was decisive for the Omnilingual lane, and that deployment-time inference code materially affected user-facing quality. We also document important limits: Track A and Track B headline numbers are not directly comparable, some development data remain non-public, and the original remote training traces were only partially retained. We present these issues explicitly because low-resource speech work benefits from transparent technical reporting rather than selective reporting.

Technical ReportTechnical ReportMay 2026

Sauti TTS Technical Report

MsingiAI

We present Sauti TTS, a Swahili text-to-speech system selected after a sequence of cloud GPU experiments covering MMS/VITSfine-tuning, a Common-Voice-based Msingi v11 acoustic vocoder pipeline, Tacotron2 artifacts, and a final F5-TTS adaptation. The report reconstructs the experiment lineage, records surviving artifacts, compares recoverable checkpoints under a common ASR probe where possible, and gives a full held-out evaluation for the selected model. The deployed system uses an F5-TTS-style fully non-autoregressive flow-matching Diffusion Transformer, Vocos waveform synthesis, Swahili normalization, reference-audio conditioning, and long-form chunking. On the complete 136-sample held-out WAXAL Swahili test split with 32 inference steps and Whisper-medium ASR, Sauti TTS obtains estimated MOS 4.485±0.176, speaker similarity 0.647, character error rate 0.1082, word error rate 0.5350, PESQ 1.196, STOI 0.006, and mean duration ratio 0.840. A new 10-sample GPU comparison probe of the recoverable MMS/VITS checkpoints gives substantially weaker ASR fidelity: CER 0.6467/WER 1.0153 for the multi-speaker 20k checkpoint and CER 0.7816/WER 1.0233 for the female-spk6 2.6k checkpoint. These results do not prove universal superiority over MMS, VITS, or Tacotron families; they justify the project-specific decision to converge on F5-TTS for Swahili because it is the only recovered path with stable reference conditioning and a completed full-test evaluation. We close with failure analysis, threats to validity, and concrete next experiments required before deployment.

Workshop PaperWorkshop PaperApril 2026

AI Across Cultures: Co-Designing Equitable and Culturally Grounded Futures

Alfred Malengo Kondoro, Jaydon Farao, Cynthia Jayne Amol, Kato Steven Mubiru, Bronson Bakunga, Gilbert Kiplangat Korir, Chris Chinenye Emezue

Artificial Intelligence (AI) and Human–Computer Interaction (HCI) are often framed as universal; yet, their foundations are shaped by the data, infrastructures, and policies of high-resource contexts. This leaves low-resource languages and culturally diverse communities underrepresented, amplifying inequities in access, adoption, and governance. This CHI 2026 workshop, AI Across Cultures: Co-Designing Equitable and Culturally Grounded Futures, brings together HCI researchers, AI/NLP practitioners, policymakers, and community leaders to reimagine AI as a sociotechnical system co-created with and for diverse contexts. We focus on aligning design methodologies, policy frameworks, and cultural philosophies to foster pluralistic, equitable, and sustainable AI ecosystems. Through case studies, participatory activities, and cross-border dialogue, the workshop will surface practices, challenges, and frameworks that integrate cultural identity, linguistic diversity, and community priorities into AI. The outcomes will include collaborative mappings, actionable guidelines, and a synthesis paper to guide future culturally grounded HCI and AI research.

PolicyPolicyJanuary 2026

AkiliX Constitution v3.0: Principles for Alignment, Safety, and East African Cultural-Linguistic Excellence

MsingiAI

The AkiliX Constitution defines the operational alignment framework for AkiliX, an East Africa–focused AI assistant. It establishes a strict priority hierarchy of safety, ethics, accuracy, helpfulness, cultural authenticity, and conversational quality. It introduces a supreme “Hard Veto” module that blocks or constrains outputs that enable physical harm, coercion, loss of dignity, high-harm misinformation, criminal facilitation, democratic manipulation, unhealthy authority substitution/dependency, or prompt-injection attempts. Beyond veto rules, the document specifies ethical principles (beneficence, non-maleficence, autonomy, justice), evidence and anti-hallucination protocols, professional boundaries for medical/legal/financial topics, and culturally grounded language behavior, including register mirroring and natural code-switching across key East African languages. It also outlines training and evaluation requirements, refusal calibration targets, multilingual quality thresholds, and a continuous improvement process to keep the system responsive to evolving regional realities such as digital access constraints and climate risk.

Position PaperPosition PaperOctober 2025

Beyond Language: Reframing LLMs in Africa Through Contextual Grounding

MsingiAI

The current discourse on artificial intelligence (AI) development for Africa predominantly emphasizes linguistic inclusion. However, this language-centric paradigm overlooks the more profound need for contextual grounding. Large Language Models (LLMs) trained primarily on Western data often misinterpret African idioms, social norms, or domain realities. As a position paper, we propose a multidimensional framework of African contextual dimensions: cultural- linguistic, socioeconomic, historical-political, and epistemological to guide the development of contextually aligned AI. Drawing on recent literature and illustrative case studies, we outline im- plications for dataset design, model evaluation, and deployment. This reframing from language coverage to contextual grounding is essential for fostering inclusive, trustworthy, and equitable AI in Africa and offers lessons for global AI ethics.