AI models are graded mostly on English. This is a mark sheet for how well frontier models actually handle Tamil — two open exams, nine models, every score reproducible with one command.
AI மாதிரிகள் பெரும்பாலும் ஆங்கிலத்திலேயே மதிப்பிடப்படுகின்றன. முன்னணி AI மாதிரிகள் தமிழை எவ்வளவு நன்றாகக் கையாளுகின்றன என்பதை அளக்கும் மதிப்பெண் அட்டவணை இது — இரண்டு திறந்த தேர்வுகள், ஒன்பது மாதிரிகள், ஒவ்வொரு மதிப்பெண்ணையும் ஒரே கட்டளையால் மீண்டும் உருவாக்கலாம்.
AI chatbots are tested on English constantly, and on Chinese, and on a handful of big European languages. Tamil — spoken by ~85 million people, official language of Tamil Nadu, Singapore and Sri Lanka — usually gets one line in a table, if it appears at all.
AI உரையாடல் மாதிரிகள் தொடர்ந்து ஆங்கிலத்திலும், சீனத்திலும், சில பெரிய ஐரோப்பிய மொழிகளிலும் சோதிக்கப்படுகின்றன. சுமார் 8.5 கோடி மக்கள் பேசும் தமிழ் — தமிழ்நாடு, சிங்கப்பூர், இலங்கை ஆகியவற்றின் அதிகாரப்பூர்வ மொழி — பெரும்பாலும் அட்டவணையில் ஒரு வரியாக மட்டுமே தோன்றுகிறது; சில நேரங்களில் அதுவும் இல்லை.
Tamil Bench fixes the measuring stick. We give AI models three exams, entirely in Tamil script, and publish the marks. Think of it as a report card: anyone can re-run the exam on any model via its API, and every answer sheet is saved so the scores can be checked.
Tamil Bench அந்த அளவுகோலைச் சரிசெய்கிறது. முழுவதும் தமிழ் எழுத்தில் அமைந்த மூன்று தேர்வுகளை AI மாதிரிகளுக்குக் கொடுத்து, மதிப்பெண்களை வெளியிடுகிறோம். இது ஒரு அறிக்கை அட்டை போன்றது: யாரும் API மூலம் எந்த மாதிரியிலும் தேர்வை மீண்டும் நடத்தலாம்; ஒவ்வொரு விடைத்தாளும் சேமிக்கப்படுவதால் மதிப்பெண்களைச் சரிபார்க்கலாம்.
Multiple-choice questions drawn from real Indian competitive exams, in Tamil. Tests whether the model knows things in Tamil — history, law, science, agriculture.
உண்மையான இந்தியப் போட்டித் தேர்வுகளில் இருந்து எடுக்கப்பட்ட தமிழ் பல்தேர்வு கேள்விகள். வரலாறு, சட்டம், அறிவியல், வேளாண்மை போன்றவற்றை மாதிரி தமிழில் அறிந்திருக்கிறதா என்பதைச் சோதிக்கிறது.
Data: MILU by AI4Bharat (with IBM Research India) — a benchmark of ~80,000 MCQs across 11 Indian languages, built from actual UPSC and state Public Service Commission papers. The Tamil paper has 6,372 questions across 41 subjects (8 domains, from Astronomy to Law); most were written for Tamil exams, not translated. We sit models for a 199-question stratified sample. License: CC-BY-4.0 (gated download on Hugging Face).
தரவு மூலம்: AI4Bharat மற்றும் IBM Research India உருவாக்கிய MILU. 11 இந்திய மொழிகளில் சுமார் 80,000 பல்தேர்வு கேள்விகள்; UPSC மற்றும் மாநில பொது சேவை ஆணையத் தேர்வுகளிலிருந்து தொகுக்கப்பட்டது. தமிழ் தொகுப்பில் 41 பாடப்பிரிவுகளில் 6,372 கேள்விகள் உள்ளன. அதிலிருந்து பாடப்பிரிவு விகிதம் காக்கப்பட்ட 199 கேள்விகளை மதிப்பிடுகிறோம். உரிமம்: CC-BY-4.0.
Read a Tamil passage, answer a question using a short phrase from it — like school reading comprehension. Tests whether the model can read and quote Tamil precisely.
ஒரு தமிழ் பத்தியைப் படித்து, அதிலிருந்து ஒரு குறுகிய சொற்றொடரைப் பயன்படுத்தி கேள்விக்கு விடையளிக்க வேண்டும் — பள்ளி வாசிப்புப் புரிதல் தேர்வு போல. தமிழ் உரையைத் துல்லியமாக படித்து மேற்கோள் காட்ட முடியுமா என்பதைச் சோதிக்கிறது.
Data: IndicQA by AI4Bharat (ACL 2023) — a hand-curated reading-comprehension dataset in 11 Indian languages. The Tamil split has 1,804 questions over 253 Wikipedia passages on Indian culture and history (1,276 answerable; 527 deliberately unanswerable to catch bluffing). We grade 100 questions per model with the official SQuAD scoring. License: CC-BY-SA-4.0.
தரவு மூலம்: AI4Bharat-ன் IndicQA. 11 இந்திய மொழிகளில் கைமுறையாகத் தொகுக்கப்பட்ட வாசிப்புப் புரிதல் தரவுத்தொகுப்பு. தமிழ் தொகுப்பில் இந்தியப் பண்பாடு மற்றும் வரலாறு குறித்த 253 Wikipedia பத்திகளில் 1,804 கேள்விகள் உள்ளன; 1,276 கேள்விகளுக்கு விடை உண்டு, 527 கேள்விகள் விடையற்றவை. அதிகாரப்பூர்வ SQuAD முறையில் ஒவ்வொரு மாதிரிக்கும் 100 கேள்விகளை மதிப்பிடுகிறோம். உரிமம்: CC-BY-SA-4.0.
| Modelமாதிரி | Accuracyசரியான விடை % | Barபட்டை | 95% CI95% நம்பிக்கை வரம்பு | |
|---|---|---|---|---|
| 1 | google/gemini-3.8-flashGemini 3.8 Flash | 94.5% | 90.4–96.9 | |
| 2 | openai/gpt-5.6-lunaGPT-5.6 Luna | 86.4% | 81.0–90.5 | |
| 3 | deepseek/deepseek-v4.1-flashDeepSeek V4.1 Flash | 80.9% | 74.9–85.8 | |
| 4 | z-ai/glm-5.3-flashGLM 5.3 Flash | 79.9% | 73.8–84.9 | |
| 5 | xiaomi/mimo-v2.5Xiaomi MiMo v2.5 | 75.9% | 69.5–81.3 | |
| 6 | qwen/qwen3.8-flashQwen 3.8 Flash | 68.8% | 62.1–74.9 | |
| 7 | inclusionai/ling-3.0-flash-vlLing 3.0 Flash VL | 64.3% | 57.5–70.6 | |
| 8 | google/gemma-4-26b-a4b-itGemma 4 26B A4B | 58.3% | 51.3–64.9 | |
| 9 | poolside/laguna-s-2.1Poolside Laguna-S 2.1 | 53.5% | 46.3–60.6 | |
| 10 | nvidia/nemotron-3.5-lightningNemotron 3.5 Lightning | 41.2% | 34.6–48.1 |
0-shot, temperature 0, letter-parse; 199 questions stratified by subject (seed 42); CI = Wilson interval. Machine-readable: results/summary.json. “Running” rows fill in as sweeps finish; parked rows await an off-peak retry.
எடுத்துக்காட்டு இல்லாத (0-shot) சோதனை; temperature 0; A/B/C/D விடை எழுத்து பகுப்பு; பாடப்பிரிவுகளின் விகிதப்படி தேர்ந்தெடுக்கப்பட்ட 199 கேள்விகள்; seed 42; CI = Wilson நம்பிக்கை இடைவெளி.
| Modelமாதிரி | Exact Matchமுழுப் பொருத்தம் | F1சொல் ஒற்றுமை | F1 BarF1 பட்டை | EM 95% CIEM நம்பிக்கை வரம்பு | |
|---|---|---|---|---|---|
| 1 | google/gemini-3.8-flashGemini 3.8 Flash | 20.0% | 42.4 | 13.3–28.9 | |
| 2 | openai/gpt-5.6-lunaGPT-5.6 Luna | 17.0% | 40.7 | 10.9–25.5 | |
| 3 | xiaomi/mimo-v2.5Xiaomi MiMo v2.5 | 20.0% | 40.1 | 13.3–28.9 | |
| 4 | google/gemma-4-26b-a4b-itGemma 4 26B A4B | 20.2% | 39.3 | 13.3–29.4 | |
| 5 | deepseek/deepseek-v4.1-flashDeepSeek V4.1 Flash | 18.0% | 38.9 | 11.7–26.7 | |
| 6 | qwen/qwen3.8-flashQwen 3.8 Flash | 22.0% | 38.1 | 15.0–31.1 | |
| 7 | inclusionai/ling-3.0-flash-vlLing 3.0 Flash VL | 14.0% | 37.4 | 8.5–22.1 | |
| 8 | z-ai/glm-5.3-flashGLM 5.3 Flash | 19.0% | 36.8 | 12.5–27.8 | |
| 9 | poolside/laguna-s-2.1Poolside Laguna-S 2.1 | 15.4% | 32.0 | 9.4–24.2 | |
| 10 | nvidia/nemotron-3.5-lightningNemotron 3.5 Lightning | 0.0% | 0.8 | 0.0–3.7 |

EM = answer matches the gold span exactly (after normalization); F1 = token overlap, 0–100. n=100 per model (seed 42).
EM = சீரமைப்புக்குப் பிறகு அதிகாரப்பூர்வ விடையுடன் முழுமையாகப் பொருந்துவது; F1 = ஒத்த சொற்களின் அளவு, 0–100. ஒவ்வொரு மாதிரிக்கும் n=100 கேள்விகள்; seed 42.
Roughly 27 of each model's 100 reading-comprehension questions were unanswerable traps — the passage simply never says. The only honest answer is “I don't know.” Bluff % is the share of traps the model answered anyway — lower is better. The verified sample answer sheet in வினா 6 shows a real one.
ஒவ்வொரு மாதிரியின் 100 வாசிப்புக் கேள்விகளில் சுமார் 27 விடையற்றப் பொய்விடைப் பொறிகள் — பத்தியில் பதிலே இல்லை. நேர்மையான பதில் “தெரியவில்லை” மட்டுமே. பொய் % என்பது மாதிரி எத்தனை பொறிகளில் பதில் சொன்னது என்பது — குறைவு சிறந்தது.
| Modelமாதிரி | Bluff % (lower = honest)பொய் % (குறைவு = நேர்மை) | Barபட்டை | Trapsபொறிகள் | |
|---|---|---|---|---|
| 1 | deepseek/deepseek-v4.1-flashDeepSeek V4.1 Flash | 48.1% | 27 traps | |
| 2 | google/gemma-4-26b-a4b-itGemma 4 26B A4B | 59.3% | 27 traps | |
| 3 | google/gemini-3.8-flashGemini 3.8 Flash | 66.7% | 27 traps | |
| 4 | qwen/qwen3.8-flashQwen 3.8 Flash | 70.4% | 27 traps | |
| 5 | xiaomi/mimo-v2.5Xiaomi MiMo v2.5 | 74.1% | 27 traps | |
| 6 | z-ai/glm-5.3-flashGLM 5.3 Flash | 74.1% | 27 traps | |
| 7 | poolside/laguna-s-2.1Poolside Laguna-S 2.1 | 85.2% | 27 traps | |
| 8 | inclusionai/ling-3.0-flash-vlLing 3.0 Flash VL | 88.9% | 27 traps | |
| 9 | nvidia/nemotron-3.5-lightningNemotron 3.5 Lightning | 88.9% | 27 traps | |
| 10 | openai/gpt-5.6-lunaGPT-5.6 Luna | 88.9% | 27 traps |
Measured on the same answer sheets as Paper 2 — no extra API calls. Abstention phrases matched in Tamil and English (தெரியவில்லை / no answer / not mentioned …). F1 on answerable questions is unaffected: traps carry empty golds, so they only ever subtract.
A three-way reading-logic exam. Given a premise and a hypothesis, the model must pick: A definitely true, B definitely false, or C cannot be decided — entirely in Tamil. A blind guess scores ~33%.
முப்பிரிவு வாசிப்பு-தர்க்கத் தேர்வு. ஒரு முன்னுரையும் ஒரு கூற்றும் கொடுக்கப்பட்டால், கூற்று A கண்டிப்பாக உண்மை, B கண்டிப்பாக தவறு, அல்லது C முடிவு செய்ய முடியாது என முடிவு செய்ய வேண்டும் — முழுவதும் தமிழில். சூதாட்டமாகவே கணித்தால் ~33% கிடைக்கும்.
| Modelமாதிரி | Accuracyசரியான விடை % | Barபட்டை | 95% CI95% நம்பிக்கை வரம்பு | |
|---|---|---|---|---|
| 1 | google/gemini-3.8-flashGemini 3.8 Flash | 75.8% | 69.3–81.2 | |
| 2 | z-ai/glm-5.3-flashGLM 5.3 Flash | 71.5% | 64.9–77.3 | |
| 3 | google/gemma-4-26b-a4b-itGemma 4 26B A4B | 64.8% | 57.8–71.2 | |
| 4 | qwen/qwen3.8-flashQwen 3.8 Flash | 63.0% | 56.1–69.4 | |
| 5 | deepseek/deepseek-v4.1-flashDeepSeek V4.1 Flash | 62.0% | 55.1–68.4 | |
| 6 | openai/gpt-5.6-lunaGPT-5.6 Luna | 61.0% | 54.1–67.5 | |
| 7 | xiaomi/mimo-v2.5Xiaomi MiMo v2.5 | 60.5% | 53.6–67.0 | |
| · | inclusionai/ling-3.0-flash-vlLing 3.0 Flash VL | parked — endpoint congested | ||
| · | nvidia/nemotron-3.5-lightningNemotron 3.5 Lightning | parked — endpoint congested | ||
| · | poolside/laguna-s-2.1Poolside Laguna-S 2.1 | running… | ||
0-shot, temperature 0, single-letter parse; n=200 per model from the 5,010-question Tamil test split (seed 42). Data: IndicXNLI (repaired) — XNLI premises/hypotheses translated to 11 Indic languages, AI4Bharat lineage. Machine-readable: results/summary.json.
Every score on this site is tallied from saved answer sheets (results/*.jsonl). Three extracts, exactly as recorded:
இந்தத் தளத்தின் ஒவ்வொரு மதிப்பெண்ணும் சேமிக்கப்பட்ட விடைத்தாள்களிலிருந்து கணக்கிடப்படுகிறது. கீழே உள்ள மூன்று எடுத்துக்காட்டுகளும் பதிவான பதில்களே.
Q: பின்வருபவர்களில் தாதா சாஹிப் பால்கே விருது 2013 பெற்ற நபர் யார்? (Who received the Dadasaheb Phalke Award for 2013?)
A. சம்பூரன் சிங் கல்ரா (Sampooran Singh Kalra — the poet Gulzar's real name) · B. விஜய ஷேசாஸ்திரி · C. ப்ரன் சிகந் · D. ரமேஷ் அகர்வால்
The model's entire reply: A — nothing else. The parser reads A, the gold is A, one mark. That is the whole contract: answer with a bare letter. When a model writes an essay instead, letter-parse digs the letter out of it — or marks the question wrong.
Passage: கோல்கொண்டா கோட்டை இந்திய தொல்லியல் துறை பட்டியலில் இடம் பெற்றுள்ளது… (The Golkonda fort features in the Archaeological Survey of India's list…)
Q: வெற்றியின் வாயில் எப்படி அழைக்கப்பட்டது? (What was the gate of victory called?) → Official answer: பதே தர்வாசா · MiMo v2.5 answered: “பதே தர்வாசா” — same words, wrapped in quotation marks. Normalization strips the punctuation, so it still scores word-for-word: full marks.
Q: அன்னை தெரேசா எப்போது தனது சேவையைத் தொடங்கினார்? (When did Mother Teresa begin her service?) → Official answer: 1948 · DeepSeek V4.1 answered: 1948 ஆம் ஆண்டில் (“in the year 1948”). The model knows the answer — but adds a grammatical flourish, so exact-match fails while F1 pays half. Best part: all nine models flunked this exact question the exact same way. Multiply it across the paper and you have the whole EM-vs-F1 gap.
இந்த முடிவுகள் நமக்குச் சொல்வது:
The whole bench is one Python file, no GPU, no fine-tuning: it calls models through an OpenAI-compatible API exactly the way a normal user would. Set OPENROUTER_API_KEY, then:
இந்த benchmark ஒரு Python கோப்பு மட்டுமே; GPU அல்லது fine-tuning தேவையில்லை. சாதாரண பயனர் பயன்படுத்துவது போல OpenAI-compatible API வழியாக மாதிரிகளை அழைக்கிறது. OPENROUTER_API_KEY-ஐ அமைத்து, கீழே உள்ள கட்டளைகளை இயக்குங்கள்: