Malayalam LLM Benchmark: Failure Analysis on Administrative Queries
An internal benchmark of leading LLMs on real-world Malayalam administrative queries reveals a ~73% failure rate — validating our regional data engine thesis.
Overview
We ran an internal benchmark of leading public LLMs — including GPT-4o, Gemini 1.5 Pro, Claude 3 Opus, and Llama 3 70B — against a curated set of 200 real-world administrative queries in Malayalam. These queries were drawn from the kinds of tasks a government office worker, a bank employee, or a primary school teacher in Kerala would actually perform.
The result: approximately 73% failure rate across all models on accurate, contextually-correct Malayalam output.
This is not a failure of intelligence. It is a failure of data. No major model has been trained on sufficient Malayalam at the density required for professional-grade outputs. English translations, transliterations, and approximate answers dominate responses where precise Malayalam is required.
What We Mean by "Failure"
We defined failure as any of the following:
- Code-switching: the model reverts to English mid-sentence for technical or administrative terms
- Transliteration errors: phonetically incorrect Malayalam rendering (Romanised Malayalam back-transliterated incorrectly)
- Semantic drift: the response is grammatically valid Malayalam but loses critical nuance from the original query (e.g., a legal deadline, a form field label, an eligibility condition)
- Hallucination of terminology: the model invents Malayalam equivalents for administrative terms that have no accepted translation
Why This Matters
Kerala has approximately 38 million Malayalam speakers. Administrative functions — land registry, ration card management, school enrollment, local governance — are conducted in Malayalam. A usable AI assistant for any of these domains requires professional-grade Malayalam output, not approximate translations.
The ~73% failure rate establishes a clear market gap. No existing model closes it. That is the gap Araskova's Malayalam-first data engine is designed to fill.
Our Approach
We are building the data moat, not training the model yet. The first phase is corpus construction: 47.5M tokens of curated Malayalam text spanning administrative documents, educational material, legal text, and conversational data — all sourced, annotated, and quality-checked by native speakers.
Once the corpus reaches sufficient density and domain coverage, fine-tuning and evaluation will follow. We will publish the benchmark methodology and dataset in full at that point.
Status
- Benchmark dataset: Complete (200 queries, 4 models, scored)
- Corpus construction: Active — 47.5M tokens and growing
- Model fine-tuning: Planned for Q4 2026
- Public paper: In preparation
This benchmark was conducted internally by Araskova Labs in June 2026. Raw data available on request.