AI in the Uzbek language is not one product but three separate jobs: understanding and writing text, transcribing voice messages, and speaking the reply out loud. They are not equally mature.
Text is the easy part for large models: Google states that all Gemini models can understand and respond in Uzbek. We found no official figures for ChatGPT: OpenAI’s open multilingual benchmark MMMLU covers 14 languages, and Uzbek is not one of them. Several cloud services offer Uzbek speech-to-text, but only some publish accuracy numbers. Fewer services than you might expect offer Uzbek text-to-speech: it is missing from Google Cloud’s voice list and from ElevenLabs’ text-to-speech models.
Text: two scripts and mixed messages
Uzbek’s official script is Latin. The law “On the Introduction of the Uzbek Alphabet Based on the Latin Script” was passed on 2 September 1993, and the current alphabet dates from the amendment of 6 May 1995: 26 letters, three digraphs (sh, ch, ng) and an apostrophe (lex.uz, in Russian). Article 2 of the same law keeps room for the Cyrillic and Arabic scripts. The full switch to Latin is being phased in; a 2021 Cabinet of Ministers resolution, No. 61, was issued for that.
In a real chat it looks like this. One customer writes “Narxi qancha?”, another “Нархи қанча?” (both mean “How much is it?”), and a third says “Салом!” and asks the rest in Russian. That gives a bot three requirements.
- Detect the language, not the alphabet. Cyrillic does not mean Russian, and Latin does not mean English.
- Reply in the same script. If the customer wrote in Cyrillic and the reply comes back in Latin, it looks like a mistake to them.
- Don’t trip over the apostrophe. People write the mark in oʻ and gʻ every which way: as ʻ, as a plain apostrophe, as a curly quote, as a backtick, or not at all. Even the text of the alphabet law on lex.uz uses a backtick, and in the professional FLORES-200 translation discussed below we counted 2,358 plain apostrophes and only 8 ʻ marks.
More variants are coming. On 10 September 2026 the Senate approved amendments to the alphabet law: the alphabet will have 28 letters and one apostrophe instead of 26 letters and three digraphs (Yangi O’zbekiston). According to Gazeta.uz, Sh, Ch, Oʻ and Gʻ are to be replaced with Ş, Ç, Ö and Ğ. The senator who presented the bill, Odiljon Mamatkarimov, argued partly from the digital side: he said more than 10 spellings of Oʻ and Gʻ are in use, so software treats the same word as different words.
The law has been sent to the president. As of 28 September there is no signed act on lex.uz, where the 1995 version is still in force. The takeaway for a bot is already clear: while the transition lasts, customers will write in at least three ways — Cyrillic, the current Latin and the new Latin. It is worth normalising them to one internal form early, before searching the catalogue and before detecting the language.
The language tax: what Uzbek costs in tokens
Models read text not letter by letter but in tokens, which are chunks of words. You pay per token, the context window is measured in tokens, and the model writes its answer one token at a time. Words the tokenizer rarely saw in training get cut into smaller pieces, so the same meaning takes more tokens.
We measured this on FLORES-200, Meta’s open parallel corpus: the same sentences from 842 web articles, translated into 200 languages. We used the devtest split, 1,012 sentences. The tokenizer is o200k_base, which according to the mapping in OpenAI’s tiktoken library is used by GPT-4o, GPT-4.1 and GPT-5 (tiktoken). As of 28 September 2026 that mapping does not say which tokenizer GPT-6 uses; if it is a different one, the ratios will differ.
| Language | Tokens for 1,012 sentences | vs English | vs Russian |
|---|---|---|---|
| English | 26,873 | ×1.00 | ×0.70 |
| Russian | 38,188 | ×1.42 | ×1.00 |
| Turkish | 38,458 | ×1.43 | ×1.01 |
| Kazakh | 42,609 | ×1.59 | ×1.12 |
| Uzbek, Latin script | 47,712 | ×1.78 | ×1.25 |
| Tajik | 50,861 | ×1.89 | ×1.33 |
| Uzbek, Cyrillic (estimate) | 52,884 | ×1.97 | ×1.38 |
What follows from this.
- The same text in Uzbek costs about a quarter more than in Russian, and nearly 40% more in Cyrillic. The price per token does not depend on the language, so the token difference goes straight into the model bill. A long Uzbek conversation also hits the context limit sooner, and replies take longer to generate.
- The newer tokenizer cut the tax sharply. On the older
cl100k_base, used by GPT-4 and GPT-3.5, the same text cost far more than English: Russian ×2.46, Uzbek Latin ×2.17, Uzbek Cyrillic ×3.55. - Grammar does not seem to be the reason. Turkish, a related Turkic language that also builds words from chains of suffixes, costs almost the same as Russian. The likely reason for the gap is how much text in each language the tokenizer saw.
- The kind of apostrophe barely matters. Replacing the plain apostrophe in oʻ and gʻ with ʻ adds less than 0.1% tokens on
o200k_base.
You can see it at word level too: “call” is 1 token, the Russian “звонок” is 3, and the Uzbek “qo’ng’iroq” and “қўнғироқ” are 5 each.
How to reproduce it. Download the FLORES-200 archive and take devtest/eng_Latn, rus_Cyrl and uzn_Latn. Install tiktoken, count len(enc.encode(line)) for every line with enc = tiktoken.get_encoding("o200k_base"), and sum. The other split, dev with 997 sentences, gives the same picture: Uzbek Latin ×1.77 against English, Russian ×1.43. FLORES has no Uzbek Cyrillic, so we transliterated uzn_Latn automatically using the letter table from the law (sh → ш, oʻ → ў, q → қ, h → ҳ and so on). Real Cyrillic text spells Russian loanwords differently, so the Cyrillic row is an estimate.
We did not measure the Gemini or Claude tokenizers; their ratios may differ.
Uzbek speech-to-text: who does it and how well
Some customers don’t type at all and record their question as a voice message. A bot that cannot transcribe speech loses those requests.
The best-known free option is OpenAI’s open Whisper model. The authors’ paper shows why it struggles with Uzbek. Its training data for Uzbek transcription was 0.3 hours, against 9,761 hours for Russian (Appendix E, Figure 11). On the FLEURS test set, Whisper large-v2 scored the following word error rates (Appendix D.2.4, Table 13): Uzbek 90.2%, Tajik 85.8%, Kazakh 37.7%, Turkish 8.4%, Russian 5.6%. WER counts wrong, missing and extra words per 100 words of the reference transcript. These figures apply to large-v2 and one dataset; newer versions are not in the paper, and the numbers should not be carried over to them. OpenAI’s own documentation puts it briefly: Whisper supports 98 languages, but accuracy varies by language.
What the other services claim in their official language lists as of September 2026:
- Google Cloud Speech-to-Text. Uzbek (uz-UZ) is available only in the Chirp models. Automatic punctuation is supported; speaker diarization is not listed for Uzbek (supported languages). That matters for call recordings.
- Gemini. Audio goes to the model directly, like text. According to the documentation, one second of audio is 32 tokens and one minute is 1,920. Per the Bot API, a Telegram voice message is OGG with the Opus codec, MP3 or M4A, and Gemini accepts all three. For audio languages, Google points to the general Gemini language list, which includes Uzbek. Audio carries no language tax: a second of Uzbek speech costs the same as a second of English. The tax applies to the transcript and the reply.
- Microsoft Azure. Uzbek in Latin script, uz-UZ, is supported for speech-to-text, including fast file transcription (language support).
- Yandex SpeechKit. The list shows “uz-UZ, Uzbek (Latin)” (recognition models, in Russian). Expect Latin-script transcripts even if the customer writes in Cyrillic.
- ElevenLabs Scribe v2. Uzbek is placed in the group with a 10–20% word error rate (documentation). That is the company’s own estimate, and the dataset is not named.
- UzbekVoice. uzbekvoice.ai offers speech-to-text and text-to-speech via API and a separate Enhanced Uzbek model for Uzbek-only audio. For Russian or mixed speech the service recommends its general model. For a bot whose customers switch between the two languages, that detail matters.
Apart from Whisper, all of this is vendors’ own claims. We found no independent 2026 comparison of these services on the same Uzbek recordings.
Uzbek text-to-speech: who can voice the reply
The choice is narrower for synthesis.
- Microsoft Azure has two voices: Madina (female) and Sardor (male).
- Yandex SpeechKit has three female voices: nigora, zamira and yulduz. Zamira and yulduz offer neutral, strict and friendly styles, and yulduz can also whisper (voice list, in Russian). There is no male Uzbek voice on the list.
- Gemini API lists Northern Uzbek in the language table of its text-to-speech models, which detect the language automatically (documentation). Northern Uzbek is the ISO 639-3 name for code uzn, the same code FLORES-200 uses for Uzbek.
- Google Cloud Text-to-Speech has no Uzbek voices in its voice list, although Russian and Turkish are there.
- ElevenLabs lists more than 70 languages for Eleven v3 on its models page, including Kazakh and Kyrgyz, but not Uzbek. The other text-to-speech models on that page don’t list it either.
- UzbekVoice offers synthesis with a choice of speaker and mood, according to its website.
Summary as of September 2026:
| Service | Speech-to-text | Text-to-speech | Watch for |
|---|---|---|---|
| Google Cloud | yes, Chirp models | not in the voice list | no speaker diarization |
| Gemini API | yes, audio as model input | yes | 32 tokens per second of audio |
| Microsoft Azure | yes, Latin script | 2 voices, female and male | — |
| Yandex SpeechKit | yes, Latin output | 3 female voices | no male voice |
| ElevenLabs | yes, Scribe v2 | not in its model lists | 10–20% WER, own estimate |
| Whisper (OpenAI) | yes, open model | not included | 90.2% WER for large-v2 on FLEURS |
| UzbekVoice | yes, dedicated Uzbek model | yes | mixed speech goes to the general model |
How to test AI in Uzbek in 30 minutes
No table can tell you how a service will handle your customers. The test takes half an hour, and you can ask any vendor to run it before you sign.
- 20 real messages from your chats. Ten in Latin script, ten in Cyrillic. Include plain apostrophes, curly ones and none at all, as in “ozbek”. Check the language and script of every reply.
- 5 voice messages. A male voice, a female voice, one recorded outdoors or in a car, one where the speaker switches from Uzbek to Russian, and one longer than a minute. Ask to see the transcript, not just the reply.
- A mixed phrase. For example, “Салом! Сколько стоит доставка?” — an Uzbek greeting followed by a Russian question. Decide in advance which language the reply should be in, then check.
- Your company name, product names and prices in som. The bot should write them exactly as given, without “translating” or rounding.
- A voice message with no speech. Ten seconds of street noise or music. The right answer is “I couldn’t make that out”, not a confident transcript. This catches models that don’t actually hear the audio and invent text.
| What you check | Sign of failure |
|---|---|
| Reply script | customer wrote in Cyrillic, reply came in Latin |
| Mixed phrase | reply in the language of the greeting, not the question |
| Names and prices | name “translated”, price rounded |
| Voice messages | voice message in Uzbek, reply in Russian |
| Noise instead of speech | the bot confidently “transcribed” something that wasn’t there |
A customer’s voice message is personal data too. What you can send to an external model after the law changed on 27 March 2026 is covered in our article on customer chats in ChatGPT and Claude.
What we don’t know
We have not verified the ElevenLabs error rates or the other vendors’ claims on our own recordings. The Whisper figures apply to large-v2 and the FLEURS set from 2022. The Uzbek Cyrillic row in the token table is an estimate based on automatic transliteration. As of 28 September there is no signed alphabet law on lex.uz, and we know the letters Ş, Ç, Ö and Ğ from press reports, not from the text of the act. Other sourced figures on AI in Uzbekistan are collected in our reference page.
Where we stand, honestly
We have no speech recognition or synthesis models of our own. We build bots from other providers’ models and test them on the client’s messages. Our main experience with Uzbek is Valli, an AI auto-responder for Telegram Business. In September 2026 we dug into why it handled Uzbek poorly and changed the following.
The language detector now recognises Uzbek in both scripts and in Latin without apostrophes, and it tells a mixed phrase like “Салом! Сколько стоит?” apart using a list of Russian words. Uzbek and Tajik messages no longer go to the small, cheap model: it garbled company names and mixed up languages. The small model only answers when the language is confidently Russian or English.
Voice messages are first transcribed by the full-size Gemini model, with Whisper turbo kept as a fallback: on Uzbek speech it gave us text that looked more like Kazakh. The lightweight Gemini version (flash-lite) did not hear the audio at all in our setup and invented a plausible transcript from the prompt. That is where step 5 of the test above comes from. We take the reply language for a voice message from the speech itself, not from the transcript.
We could not get the model to answer in Cyrillic with a prompt alone: it kept drifting back to Latin. Code fixed it. If the customer wrote in Cyrillic and the reply came back in Latin, a separate step rewrites it in Cyrillic and keeps brand names in Latin.
These are observations from our own conversations, not measurements. We have no percentages, so there are none in this article.
How we build an AI agent that replies in Russian and in Uzbek in both scripts, and test its answers on your real messages before launch, is described on the AI agents for business page. If you need a simpler bot that takes requests, answers routine questions and hands everything else to a person, see business chatbots.
If you want your bot or vendor put through this test, describe the task in the questionnaire on our home page or message us on Telegram. The breakdown of one process takes 48 hours and is free.