Twi and Ghanaian English speech synthesis, driven by IPA phonemes. Code · Model
Akwaaba, wo ho te sɛn? Me da wo ase paa.
| Voice | Sample | Twi / code-switch error |
|---|---|---|
| twi-1 | 34% / 60% | |
| twi-2 | 29% / 61% | |
| twi-3 | 29% / 61% | |
| twi-4 | 32% / 61% | |
| twi-5 | 28% / 61% | |
| twi-6 | 27% / 61% | |
| twi-7 | 27% / 61% | |
| twi-8 | 30% / 61% | |
| twi-9 | 28% / 66% | |
| twi-10 | 30% / 64% | |
| twi-11 | 30% / 63% | |
| twi-12 | 30% / 64% |
Mepɛ sɛ mesua [computer science] wɔ [University of Ghana].
| Voice | Sample | Twi / code-switch error |
|---|---|---|
| twi-1 | 34% / 60% | |
| twi-2 | 29% / 61% | |
| twi-3 | 29% / 61% | |
| twi-4 | 32% / 61% | |
| twi-5 | 28% / 61% | |
| twi-6 | 27% / 61% | |
| twi-7 | 27% / 61% | |
| twi-8 | 30% / 61% | |
| twi-9 | 28% / 66% | |
| twi-10 | 30% / 64% | |
| twi-11 | 30% / 63% | |
| twi-12 | 30% / 64% |
Different sentence types, using the best voice for each mode.
| Kind | Text | Sample |
|---|---|---|
| greeting | Akwaaba! Yɛma wo akwaaba wɔ Ghana. | |
| statement | Ghana yɛ ɔman a ɛwɔ Afrika atɔeɛ fam. | |
| question | Wo din de sɛn? Wofiri he na woreba? | |
| long | Anɔpa yi, ɔsoro abue na awia bɔ. Nnipa pii firi wɔn afie mu rekɔ adwuma, na mmɔfra nso rekɔ sukuu. | |
| numbers | Yɛn nsa kaa nnipa apem ne ahanum wɔ ɔmantam no mu. | |
| news | Ɔkyerɛkyerɛni no kaa sɛ [the examination] bɛba [next week]. | |
| institution | [Bank of Ghana] abɔ [interest rate] no so bio. | |
| english | Good morning, and welcome to the news. |
Voices are pseudo-speakers — derived by clustering x-vectors over unlabelled broadcast audio, because neither source corpus had speaker labels. One real person may appear as two voices, and no voice is a consented identity.
They are ranked by measured intelligibility, not training hours, which turned out
to predict almost nothing: twi-1 is the best code-switch voice yet 21st of 30 on pure
Twi, and two of the three best Twi voices have under 3.3 h of audio each. Use
tiers.codeswitch for text mixing English into Twi and tiers.twi_only for
pure Twi — the two rankings disagree sharply.
English is audibly weaker than Twi and band-limited to 8 kHz, because the English training audio was 16 kHz where the Twi was 24 kHz. Round-trip phoneme error is 33.5% for Twi against a 25.9% floor, and 59.5% for English against 32.2%.