Measuring what AI models can actually do in Bantu languages, one capability at a time.
Each component measures one capability, has its own board and its own status. Families group them; new components join as their ground truth matures.
the building blocks a language is made of
Can a model list every building block its language is made of? English has 26 letters; Bemba has 480 syllables.
counting, and everything built on it
Can a model operate the number system, or only recall number words? Money, dates and sums are TRACKS inside it, not separate components.
the machinery that makes a sentence agree with itself; every component here is a dimension group of the concord matrix
The family is declared; a scored component joins once its ground truth is ready.
what people actually say to each other
The family is declared; a scored component joins once its ground truth is ready.
capability where it has consequences
Can a model name the body, understand what a patient says, and say it back? And what does it invent when it cannot — on danger signs?
The same principles hold for every component.
Every answer key comes from native speakers, and every attested alternative form is accepted.
A component measures a single capability, such as the alphabet, numbers or clinical language, so a score says exactly what was tested.
Answer keys are never published. Models are scored blind, and editions are frozen so results stay comparable.
Components grow across more languages, more items per language, and more of the standard, measured against 459 released inventories.