How do LLMs understand Gujarati?

One of my favorite ways of testing LLM-powered apps at Jeavio is to ask them questions in transliterated Hindi or Gujarati. I ask questions in Latin script and see how the application responds.

When building chat apps, we are often given instructions by our clients that the bot should only support English. This is an interesting test case on the type of guardrails that our engineers have built into the app.

The more interesting point is why models behave this way 🤔.

Take the screenshot below. Here I ask ChatGPT a question in phonetic Gujarati about Horza, a character in Iain M. Banks’s “Consider Phlebas.” The model understood and responded in phonetic Gujarati. The style was very formal and not quite like how most people speak, but it was recognizable as Gujarati.

Intuitively, you would assume that models are trained on multilingual data and can respond to questions in multiple languages. Gujarati training data in -> Gujarati output out.

However, it is unlikely that a significant amount of Gujarati language analysis of an Iain M. Banks book is available.

There is some interesting research in this space (citations below):

Shared semantic spaces across languages
It appears that LLMs learn a shared semantic space, allowing them to take content from a high-resource language, such as English (with lots of nerdy sci-fi commentary), and respond in a lower-resource language, like Gujarati. The model somehow maps concepts across languages even when direct translations don’t exist in the training data.

Transliteration without explicit training
Gujarati has its own script, of course, so how do models understand transliterated languages? While tokenizers are typically biased toward their training distribution, models appear to learn mappings between transliterated tokens and semantic concepts, despite not being explicitly designed for this purpose. The model figures out that “Horza” in Latin-script Gujarati refers to the same entity as “Horza” in English.

Cross-lingual knowledge alignment
Research also shows that the internal knowledge representations seem to align across languages. This enables translation between language pairs that lack a shared vocabulary. The model builds bridges where none existed before.

Emergent, not designed
This behavior is all emergent! Models weren’t explicitly trained on transliteration pairs or given instructions to handle Latin-script versions of non-Latin languages. They figured it out on their own! Smart models 🧐.

So LLMs are weird.

While we learn more about them, some of their behaviors are still emergent and unpredictable. This makes evaluations extremely important when building with LLMs – something that my team at Jeavio is learning very quickly.

Meanwhile, my patient and generous QA teams continue to tolerate my weird edge cases involving transliterated languages and decades-old science fiction.


Citations (sourced via ChatGPT’s Deep Research mode)

  1. Language Models are Unsupervised Multitask Learners (OpenAI – 2019)
  2. ByT5: Towards a Token-Free Future with Pre-trained Byte-to-Byte Models (Google Research – 2022)
  3. Crosslingual Generalization through Multitask Finetuning (2023)

Here’s the output from DeepResearch summarizing the results:


Modern multilingual LLMs don’t keep a single “language-free knowledge graph,” but they do learn a shared semantic space that lets them read transliterated inputs (e.g., Latin-script Gujarati), retrieve facts learned mainly from high-resource languages like English, and answer back in the user’s language. This works because their tokenization and modeling are language-agnostic: subword/byte tokenizers (e.g., SentencePiece) and byte-level models (e.g., ByT5) reliably parse mixed or nonstandard text and are robust to spelling/romanization noise.  During pretraining on many languages, the model’s internal representations align across languages—even without a shared vocabulary—so knowledge learned in one language can be accessed from another.  Evidence from multilingual machine translation shows similar “interlingua-like” behavior via zero-shot translation between unseen language pairs.  Probing studies further show that models can recall factual knowledge across languages, though performance varies by language and prompt.  Public resources also provide training/evaluation signal for romanized/code-mixed text (e.g., the Dakshina dataset for Indic languages), which reinforces these abilities.  The caveat: cross-lingual answers are not perfectly consistent—recent analyses find significant variability in factual consistency across language pairs—so quality can be uneven, especially for low-resource languages.