BantuNomics builds the missing foundation that makes AI actually work for the 400 million people who speak Bantu languages — starting from the most basic building block there is: the complete set of syllables each language is made of.
| ∅ | a | e | i | o | u |
| b | ba | be | bi | bo | bu |
| m | ma | me | mi | mo | mu |
| k | ka | ke | ki | ko | ku |
| ng' | ng'a | ng'e | ng'i | ng'o | ng'u |
| mb | mba | mbe | mbi | mbo | mbu |
Swahili, Zulu, Shona, Bemba, Nyanja and roughly 500 others — the Bantu family — are spoken across sub-Saharan Africa by more than 400 million people. That's more than the entire population of the United States. Yet ask the best AI on Earth something in most of those languages and it falls on its face: fluent-sounding, confidently wrong.
Why? Today's AI learned language by reading the internet. The internet barely contains these languages — so the model never really learned them; it guesses. And the text that exists throws away the most important part: in many of these languages the pitch of your voice changes a word's meaning, and flat text is silent about melody — like sheet music printed as lyrics.
The result isn't an inconvenience — it's a wall between AI and hundreds of millions of people in healthcare, education, and finance, where a mistranslation is dangerous and "close enough" is not enough.
English has 26 letters; every English word is built from them. Bantu languages are built from syllables — ba be bi bo bu — and a child learns them one at a time. The full set of a language's legal syllables is its operating alphabet. It is closed, finite, and knowable: you can write down every one. Bemba has about 480.
BantuNomics has established and verified that complete operating alphabet for each language — the real set of legal syllables, worked out from the language's own sound system and checked against ground truth. Every syllable fits a shared structural template σ → (N)(C)(H)(G)V, which lets the inventory be laid out as one clean grid — a first in Bantu history.
Here's what makes it valuable: no AI on Earth can produce it unaided. Getting a model to guess at syllables is easy; knowing which ones are actually real, and that the set is complete, is the hard part — and that certainty comes only from native ground truth. This isn't data you could scrape harder — it's a foundational fact about each language. Think periodic table, not spreadsheet. It's infrastructure, not a dataset.
Anyone can generate a list of syllables. Only native ground truth can tell you which ones are real — and that is the part no model can shortcut.
We turned that into a public benchmark, L26. In English, today's leading models are flawless. On a Bantu language the best model manages roughly half the score, and most collapse far below. The gap is real, measurable, and reproducible — check it in a minute at l26.ai.
Nobody buys "a platform." You arrive caring about one thing — clinical speech, tone, translation, tokenization. So each domain is its own product, grounded in published Bantu scholarship and consented native audio, solving a concrete problem for a specific buyer. FSI is the foundational domain everything decomposes back into — but it is one product among several. Open the one that's yours.
Consented patient audio that code-switches into English to name symptoms yields, from a single recording, clinical-term ground truth, code-switch ASR, an accented-English fairness benchmark, native clinical TTS, and a self-grading eval against frontier models — anchored to an interactive body figure. Frontier ASR still mishears "heartbeat." This is how health AI gets actually safe in Bantu.
The closed, verified set of sounds each language is built from, certified against native ground truth and laid out as one clean grid. Generation is cheap; knowing which syllables are real, and that the set is complete, is the part no model can shortcut. The base layer every other domain — and every model — stands on.
In Bantu languages pitch and vowel length carry meaning: one written word can be a different word entirely. We prove it with minimal homograph pairs and restore it in consented native audio — the signal flat text silently throws away.
Explore Tone →Not scattered per-language wordlists — an aligned matrix anchored to a shared English reference, so a model learns the whole family at once. Each entry carries two-mode native audio: clean Bantu and the code-switch people really use.
Explore Nouns →A Bantu verb expands into hundreds of forms through ordered building blocks. We ship the generative engine and the corpus together, so a model sees the system that produces the forms — not just scattered samples of them.
Explore Verbs →Every numeral rendered and voiced two ways: pure Bantu, and the everyday Bantu-English mix real speakers use for money, dates, and quantities. Consented native voice, ready for TTS and ASR.
Explore Numbers →Bantu concord threads agreement through every word that follows the noun. We certify it cell by cell, so a model can build sentences that are actually grammatical — not fluent-looking word salad that a native speaker would never say.
Explore Grammar →The same connected speech as clean native Bantu, natural Bantu-English code-switch, and accented English — time-aligned across all three. The code-switch and accented-speech layers monolingual corpora simply don't have.
Explore Stories →Every domain is worked out from a language's real grammar and sound system and verified against ground truth — the reason no model can reproduce it, and the reason it doesn't drift or go stale.
Concepts are defined once and anchored to a common English reference, then filled in per language — an aligned multilingual matrix, not disconnected wordlists.
Audio is recorded by native speakers who are compensated and give explicit consent, via the amina.ai "Our Words" platform. Consent is checked on every use; identities never exposed.
Each consented recording does the work of several separate datasets a buyer would otherwise piece together from many sources. License it once; get many kinds of value.
A system that calls itself general intelligence but can't handle the foundational layer of any language spoken by hundreds of millions of people has a real hole in the map. We fill it — as licensable, aligned infrastructure that plugs into your stack via API and standard tooling.
This is the groundwork that lets AI serve people it currently leaves behind — built ethically, from the community's own voices, as public-good infrastructure other tools can stand on.
Four ways in, each a real step — no obligation to climb them in order. The paid pilot is 100% creditable toward a subscription, so nothing you spend to evaluate is lost.
See the whole thesis, run the public alphabet test, read the L26 benchmark. Verify the gap yourself — no account.
Try it free →Score your own models on the Alphabet Test for the assigned languages (Bemba + Nyanja) — rates + L26, saved. The inventory itself begins at Pilot.
Request access → Most teams start hereA scoped, paid proof on your own held-out data, across your chosen languages and every domain. 100% credited toward a subscription.
Scope a pilot →License the whole living platform — every domain, every released language, the full consented corpus, and everything added while you're subscribed.
Talk to us →A separate, non-commercial track for grant, philanthropic, and impact funders — a curated Showroom view of the work and where it's going. Not a subscription, not for AI labs.
BantuNomics is a living platform, not a finished dataset — and it grows in two directions at once: wider, as we add more of the 500-plus Bantu languages, and deeper, as each language gains more words, recordings, pronunciation, and grammar. Languages are alive, so this growth never stops — there is no finish line, and everything is continuously curated.
Today it covers a share of the family — the foundation reaches furthest, while the layers above and the consented audio are earlier and expanding all the time. That is precisely why the top tier is a subscription to the living platform, not a purchase of a finished corpus: you get everything that exists today and everything added while you're subscribed.
The live leaderboard — watch today's frontier models ace English and fail the same alphabet test on a Bantu language.
Give any AI the same task we give ours, in about a minute. See the gap first-hand, no account needed.
Start a free evaluation, scope a paid pilot on your own data, or open a subscription conversation.