◀ Knowledge hub

03, AI & intelligence

Multilingual voice, including Swahili and Sheng

codeAmani Labs Engineering
Cinematic still for Multilingual voice, including Swahili and Sheng

Voice is the interface for people text leaves behind

A lot of product thinking assumes the user is comfortable reading and typing in English on a screen. For a large share of users that assumption is wrong, and voice is not a luxury feature, it is the accessible path. Text to speech and speech to text in the user's actual language, including Swahili and Sheng, opens the product to people a text only build quietly excludes.

Two directions, two jobs

Text to speech reads your content aloud: an interactive voice response menu, a spoken confirmation, a voice note summary. Speech to text turns what the user says into something your system can act on: a spoken query, a dictated message, a voice driven form. Most voice features are a loop of the two, listen, understand, respond.

// synthesize a spoken confirmation in the user's language
const audio = await elevenlabs.textToSpeech.convert({
  voiceId: SWAHILI_VOICE,
  text: "Malipo yamekamilika. Asante.",
  modelId: "multilingual",
});

Language coverage is the whole point

A multilingual voice model that handles Swahili well, and can cope with the code switching of Sheng, is doing something the default English voices cannot. In a market where the lingua franca of the street is not the lingua franca of the software, matching the user's language is the difference between a feature people use and a feature people abandon.

Wire it to the channels that reach everyone

Voice is most powerful when paired with the communications layer. A voice response system reachable over a normal phone call serves users with no smartphone and no data, the same constituency the USSD and SMS channels serve. The model produces the audio; the telephony channel delivers it to a basic handset. Together they reach people that an app store download never will.

Keep it natural

The fastest way to make voice feel cheap is robotic phrasing and awkward pauses. Write the spoken copy the way a person would say it, not the way a form would print it, and test it by ear with native speakers. Voice that sounds human gets trusted; voice that sounds like a machine reading a spreadsheet gets hung up on.

Qualified conversation

Have a build to de-risk? Let's talk.

Tell us what you are building. We respond within two business days.