ThamizhKanimai தமிழ்க் கணிமை
ThamizhKanimai · தமிழ்க் கணிமை Benchmark Series · Sept 2026

தமிழ்த் திறனாய்வுTamil Bench

AI models are graded mostly on English. This is a mark sheet for how well frontier models actually handle Tamil — two open exams, nine models, every score reproducible with one command.

AI மாதிரிகள் பெரும்பாலும் ஆங்கிலத்திலேயே மதிப்பிடப்படுகின்றன. முன்னணி AI மாதிரிகள் தமிழை எவ்வளவு நன்றாகக் கையாளுகின்றன என்பதை அளக்கும் மதிப்பெண் அட்டவணை இது — இரண்டு திறந்த தேர்வுகள், ஒன்பது மாதிரிகள், ஒவ்வொரு மதிப்பெண்ணையும் ஒரே கட்டளையால் மீண்டும் உருவாக்கலாம்.

வினா 1

What is Tamil Bench? (ELI5)தமிழ்த் திறனாய்வு என்றால் என்ன? (எளிய விளக்கம்)

[10 marks]

AI chatbots are tested on English constantly, and on Chinese, and on a handful of big European languages. Tamil — spoken by ~85 million people, official language of Tamil Nadu, Singapore and Sri Lanka — usually gets one line in a table, if it appears at all.

AI உரையாடல் மாதிரிகள் தொடர்ந்து ஆங்கிலத்திலும், சீனத்திலும், சில பெரிய ஐரோப்பிய மொழிகளிலும் சோதிக்கப்படுகின்றன. சுமார் 8.5 கோடி மக்கள் பேசும் தமிழ் — தமிழ்நாடு, சிங்கப்பூர், இலங்கை ஆகியவற்றின் அதிகாரப்பூர்வ மொழி — பெரும்பாலும் அட்டவணையில் ஒரு வரியாக மட்டுமே தோன்றுகிறது; சில நேரங்களில் அதுவும் இல்லை.

Tamil Bench fixes the measuring stick. We give AI models three exams, entirely in Tamil script, and publish the marks. Think of it as a report card: anyone can re-run the exam on any model via its API, and every answer sheet is saved so the scores can be checked.

Tamil Bench அந்த அளவுகோலைச் சரிசெய்கிறது. முழுவதும் தமிழ் எழுத்தில் அமைந்த மூன்று தேர்வுகளை AI மாதிரிகளுக்குக் கொடுத்து, மதிப்பெண்களை வெளியிடுகிறோம். இது ஒரு அறிக்கை அட்டை போன்றது: யாரும் API மூலம் எந்த மாதிரியிலும் தேர்வை மீண்டும் நடத்தலாம்; ஒவ்வொரு விடைத்தாளும் சேமிக்கப்படுவதால் மதிப்பெண்களைச் சரிபார்க்கலாம்.

If you know nothing about AI AI பற்றி ஒன்றும் தெரியவில்லை என்றால் Imagine nine AI models — different systems with the same claim: “we know Tamil.” We seat them for two papers: a multiple-choice general-knowledge exam, and a reading-comprehension test where they must underline the exact phrase in a passage that answers a question. Then we tally the marks and put the topper on the board. That's this website. தமிழ் தெரியும் என்று கூறும் ஒன்பது AI மாதிரிகளை நினைத்துப் பாருங்கள். அவற்றுக்கு இரண்டு தாள்கள் கொடுக்கப்படுகின்றன: பல்தேர்வு பொது அறிவுத் தேர்வு ஒன்று; ஒரு பத்தியைப் படித்து, கேள்விக்கான சரியான சொற்றொடரை அதிலிருந்து எடுத்துக் கூற வேண்டிய வாசிப்புப் புரிதல் தேர்வு ஒன்று. பிறகு மதிப்பெண்களை எண்ணி, முதலிடத்தைப் பலகையில் காட்டுகிறோம். அதுதான் இந்த இணையதளம்.

MILU — the MCQ exam

பொது அறிவுத் தேர்வு

Multiple-choice questions drawn from real Indian competitive exams, in Tamil. Tests whether the model knows things in Tamil — history, law, science, agriculture.

உண்மையான இந்தியப் போட்டித் தேர்வுகளில் இருந்து எடுக்கப்பட்ட தமிழ் பல்தேர்வு கேள்விகள். வரலாறு, சட்டம், அறிவியல், வேளாண்மை போன்றவற்றை மாதிரி தமிழில் அறிந்திருக்கிறதா என்பதைச் சோதிக்கிறது.

Data: MILU by AI4Bharat (with IBM Research India) — a benchmark of ~80,000 MCQs across 11 Indian languages, built from actual UPSC and state Public Service Commission papers. The Tamil paper has 6,372 questions across 41 subjects (8 domains, from Astronomy to Law); most were written for Tamil exams, not translated. We sit models for a 199-question stratified sample. License: CC-BY-4.0 (gated download on Hugging Face).

தரவு மூலம்: AI4Bharat மற்றும் IBM Research India உருவாக்கிய MILU. 11 இந்திய மொழிகளில் சுமார் 80,000 பல்தேர்வு கேள்விகள்; UPSC மற்றும் மாநில பொது சேவை ஆணையத் தேர்வுகளிலிருந்து தொகுக்கப்பட்டது. தமிழ் தொகுப்பில் 41 பாடப்பிரிவுகளில் 6,372 கேள்விகள் உள்ளன. அதிலிருந்து பாடப்பிரிவு விகிதம் காக்கப்பட்ட 199 கேள்விகளை மதிப்பிடுகிறோம். உரிமம்: CC-BY-4.0.

Score = % of questions answered correctly, with 95% confidence intervalsமதிப்பெண் = சரியாக விடையளிக்கப்பட்ட கேள்விகளின் சதவீதம்; 95% நம்பிக்கை இடைவெளியுடன்

IndicQA — the comprehension exam

வாசிப்புப் புரிதல் தேர்வு

Read a Tamil passage, answer a question using a short phrase from it — like school reading comprehension. Tests whether the model can read and quote Tamil precisely.

ஒரு தமிழ் பத்தியைப் படித்து, அதிலிருந்து ஒரு குறுகிய சொற்றொடரைப் பயன்படுத்தி கேள்விக்கு விடையளிக்க வேண்டும் — பள்ளி வாசிப்புப் புரிதல் தேர்வு போல. தமிழ் உரையைத் துல்லியமாக படித்து மேற்கோள் காட்ட முடியுமா என்பதைச் சோதிக்கிறது.

Data: IndicQA by AI4Bharat (ACL 2023) — a hand-curated reading-comprehension dataset in 11 Indian languages. The Tamil split has 1,804 questions over 253 Wikipedia passages on Indian culture and history (1,276 answerable; 527 deliberately unanswerable to catch bluffing). We grade 100 questions per model with the official SQuAD scoring. License: CC-BY-SA-4.0.

தரவு மூலம்: AI4Bharat-ன் IndicQA. 11 இந்திய மொழிகளில் கைமுறையாகத் தொகுக்கப்பட்ட வாசிப்புப் புரிதல் தரவுத்தொகுப்பு. தமிழ் தொகுப்பில் இந்தியப் பண்பாடு மற்றும் வரலாறு குறித்த 253 Wikipedia பத்திகளில் 1,804 கேள்விகள் உள்ளன; 1,276 கேள்விகளுக்கு விடை உண்டு, 527 கேள்விகள் விடையற்றவை. அதிகாரப்பூர்வ SQuAD முறையில் ஒவ்வொரு மாதிரிக்கும் 100 கேள்விகளை மதிப்பிடுகிறோம். உரிமம்: CC-BY-SA-4.0.

Score = Exact Match (did you quote the exact phrase?) and F1 (partial credit for overlapping words)மதிப்பெண் = Exact Match (அதே சொற்றொடரா?) மற்றும் F1 (ஒத்த சொற்களுக்கு பகுதி மதிப்பெண்)
வினா 2

Scoreboard · Paper 1: MILUமதிப்பெண் அட்டவணை · தாள் 1: MILU

[199 marks]
ModelமாதிரிAccuracyசரியான விடை %Barபட்டை95% CI95% நம்பிக்கை வரம்பு
1google/gemini-3.8-flashGemini 3.8 Flash94.5%90.4–96.9
2openai/gpt-5.6-lunaGPT-5.6 Luna86.4%81.0–90.5
3deepseek/deepseek-v4.1-flashDeepSeek V4.1 Flash80.9%74.9–85.8
4z-ai/glm-5.3-flashGLM 5.3 Flash79.9%73.8–84.9
5xiaomi/mimo-v2.5Xiaomi MiMo v2.575.9%69.5–81.3
6qwen/qwen3.8-flashQwen 3.8 Flash68.8%62.1–74.9
7inclusionai/ling-3.0-flash-vlLing 3.0 Flash VL64.3%57.5–70.6
8google/gemma-4-26b-a4b-itGemma 4 26B A4B58.3%51.3–64.9
9poolside/laguna-s-2.1Poolside Laguna-S 2.153.5%46.3–60.6
10nvidia/nemotron-3.5-lightningNemotron 3.5 Lightning41.2%34.6–48.1

0-shot, temperature 0, letter-parse; 199 questions stratified by subject (seed 42); CI = Wilson interval. Machine-readable: results/summary.json. “Running” rows fill in as sweeps finish; parked rows await an off-peak retry.

எடுத்துக்காட்டு இல்லாத (0-shot) சோதனை; temperature 0; A/B/C/D விடை எழுத்து பகுப்பு; பாடப்பிரிவுகளின் விகிதப்படி தேர்ந்தெடுக்கப்பட்ட 199 கேள்விகள்; seed 42; CI = Wilson நம்பிக்கை இடைவெளி.

வினா 3

Scoreboard · Paper 2: IndicQAமதிப்பெண் அட்டவணை · தாள் 2: IndicQA

[100 marks]
ModelமாதிரிExact Matchமுழுப் பொருத்தம்F1சொல் ஒற்றுமைF1 BarF1 பட்டைEM 95% CIEM நம்பிக்கை வரம்பு
1google/gemini-3.8-flashGemini 3.8 Flash20.0%42.413.3–28.9
2openai/gpt-5.6-lunaGPT-5.6 Luna17.0%40.710.9–25.5
3xiaomi/mimo-v2.5Xiaomi MiMo v2.520.0%40.113.3–28.9
4google/gemma-4-26b-a4b-itGemma 4 26B A4B20.2%39.313.3–29.4
5deepseek/deepseek-v4.1-flashDeepSeek V4.1 Flash18.0%38.911.7–26.7
6qwen/qwen3.8-flashQwen 3.8 Flash22.0%38.115.0–31.1
7inclusionai/ling-3.0-flash-vlLing 3.0 Flash VL14.0%37.48.5–22.1
8z-ai/glm-5.3-flashGLM 5.3 Flash19.0%36.812.5–27.8
9poolside/laguna-s-2.1Poolside Laguna-S 2.115.4%32.09.4–24.2
10nvidia/nemotron-3.5-lightningNemotron 3.5 Lightning0.0%0.80.0–3.7
Tamil Bench comparison chart: 9 models across MILU accuracy, IndicQA exact match, F1 and IndicXNLI accuracy
Report-card image · all metrics side by side — share freely (CC-BY-SA, attribute tamil-bench).

EM = answer matches the gold span exactly (after normalization); F1 = token overlap, 0–100. n=100 per model (seed 42).

EM = சீரமைப்புக்குப் பிறகு அதிகாரப்பூர்வ விடையுடன் முழுமையாகப் பொருந்துவது; F1 = ஒத்த சொற்களின் அளவு, 0–100. ஒவ்வொரு மாதிரிக்கும் n=100 கேள்விகள்; seed 42.

வினா 4

Scoreboard · Paper 2b: the bluff catchமதிப்பெண் அட்டணை · தாள் 2b: பொய்விடைப் பிடிக்கும் பந்தயம்

[traps]

Roughly 27 of each model's 100 reading-comprehension questions were unanswerable traps — the passage simply never says. The only honest answer is “I don't know.” Bluff % is the share of traps the model answered anyway — lower is better. The verified sample answer sheet in வினா 6 shows a real one.

ஒவ்வொரு மாதிரியின் 100 வாசிப்புக் கேள்விகளில் சுமார் 27 விடையற்றப் பொய்விடைப் பொறிகள் — பத்தியில் பதிலே இல்லை. நேர்மையான பதில் “தெரியவில்லை” மட்டுமே. பொய் % என்பது மாதிரி எத்தனை பொறிகளில் பதில் சொன்னது என்பது — குறைவு சிறந்தது.

ModelமாதிரிBluff % (lower = honest)பொய் % (குறைவு = நேர்மை)Barபட்டைTrapsபொறிகள்
1deepseek/deepseek-v4.1-flashDeepSeek V4.1 Flash48.1%27 traps
2google/gemma-4-26b-a4b-itGemma 4 26B A4B59.3%27 traps
3google/gemini-3.8-flashGemini 3.8 Flash66.7%27 traps
4qwen/qwen3.8-flashQwen 3.8 Flash70.4%27 traps
5xiaomi/mimo-v2.5Xiaomi MiMo v2.574.1%27 traps
6z-ai/glm-5.3-flashGLM 5.3 Flash74.1%27 traps
7poolside/laguna-s-2.1Poolside Laguna-S 2.185.2%27 traps
8inclusionai/ling-3.0-flash-vlLing 3.0 Flash VL88.9%27 traps
9nvidia/nemotron-3.5-lightningNemotron 3.5 Lightning88.9%27 traps
10openai/gpt-5.6-lunaGPT-5.6 Luna88.9%27 traps

Measured on the same answer sheets as Paper 2 — no extra API calls. Abstention phrases matched in Tamil and English (தெரியவில்லை / no answer / not mentioned …). F1 on answerable questions is unaffected: traps carry empty golds, so they only ever subtract.

வினா 5

Scoreboard · Paper 3: IndicXNLIமதிப்பெண் அட்டணை · தாள் 3: IndicXNLI

[200 marks]

A three-way reading-logic exam. Given a premise and a hypothesis, the model must pick: A definitely true, B definitely false, or C cannot be decided — entirely in Tamil. A blind guess scores ~33%.

முப்பிரிவு வாசிப்பு-தர்க்கத் தேர்வு. ஒரு முன்னுரையும் ஒரு கூற்றும் கொடுக்கப்பட்டால், கூற்று A கண்டிப்பாக உண்மை, B கண்டிப்பாக தவறு, அல்லது C முடிவு செய்ய முடியாது என முடிவு செய்ய வேண்டும் — முழுவதும் தமிழில். சூதாட்டமாகவே கணித்தால் ~33% கிடைக்கும்.

ModelமாதிரிAccuracyசரியான விடை %Barபட்டை95% CI95% நம்பிக்கை வரம்பு
1google/gemini-3.8-flashGemini 3.8 Flash75.8%69.3–81.2
2z-ai/glm-5.3-flashGLM 5.3 Flash71.5%64.9–77.3
3google/gemma-4-26b-a4b-itGemma 4 26B A4B64.8%57.8–71.2
4qwen/qwen3.8-flashQwen 3.8 Flash63.0%56.1–69.4
5deepseek/deepseek-v4.1-flashDeepSeek V4.1 Flash62.0%55.1–68.4
6openai/gpt-5.6-lunaGPT-5.6 Luna61.0%54.1–67.5
7xiaomi/mimo-v2.5Xiaomi MiMo v2.560.5%53.6–67.0
·inclusionai/ling-3.0-flash-vlLing 3.0 Flash VLparked — endpoint congested
·nvidia/nemotron-3.5-lightningNemotron 3.5 Lightningparked — endpoint congested
·poolside/laguna-s-2.1Poolside Laguna-S 2.1running…

0-shot, temperature 0, single-letter parse; n=200 per model from the 5,010-question Tamil test split (seed 42). Data: IndicXNLI (repaired) — XNLI premises/hypotheses translated to 11 Indic languages, AI4Bharat lineage. Machine-readable: results/summary.json.

வினா 6

How to read the marks — rubrics & jargon in plain languageமதிப்பெண்களை எப்படி படிப்பது — அளவுகோல்கள் மற்றும் சொற்கள்

[glossary]
AccuracyThe share of exam questions answered correctly. 94.5% means roughly 188 of 199 questions right.தேர்வு கேள்விகளில் சரியாக விடையளிக்கப்பட்டவற்றின் பங்கு. 94.5% என்றால் 199 கேள்விகளில் சுமார் 188 சரி.
EM · Exact MatchFull marks only if the answer is word-for-word the official answer (ignoring capital letters, punctuation and extra spaces). Harsh but unambiguous.அதிகாரப்பூர்வ விடையுடன் சொல்-சொல்லாகப் பொருந்தினால் மட்டுமே முழு மதிப்பெண். நிறுத்தற்குறி, இடைவெளி போன்றவை சீரமைக்கப்பட்ட பிறகு ஒப்பிடப்படும்.
F1Partial credit: the share of words the model’s answer shares with the official answer. 100 = same words, 0 = no overlap. It catches “almost right” answers that EM fails.பகுதி மதிப்பெண்: மாதிரியின் விடையும் அதிகாரப்பூர்வ விடையும் பகிர்ந்து கொள்ளும் சொற்களின் அளவு. 100 = முழுப் பொருத்தம்; 0 = ஒற்றுமை இல்லை. “கிட்டத்தட்ட சரி” விடைகளையும் இது காட்டும்.
95% CIWe grade a sample of questions, not the whole set. The CI is the range where the model’s true score would land in 95 of 100 re-samples. Non-overlapping ranges = a real gap, not luck.முழுத் தொகுப்பையும் அல்ல, ஒரு மாதிரியையே மதிப்பிடுகிறோம். மீண்டும் 100 முறை மாதிரி எடுத்தால் 95 முறை உண்மையான மதிப்பெண் விழும் வரம்பே CI. வரம்புகள் ஒட்டாதபோது வித்தியாசம் நம்பத்தகுந்தது.
0-shotThe model gets no worked examples before the exam — just the question, cold.தேர்வுக்கு முன் மாதிரிக்கு எடுத்துக்காட்டு விடைகள் எதுவும் கொடுக்கப்படாது — கேள்வி மட்டும் நேரடியாக வழங்கப்படும்.
temperature 0The model’s randomness dial is set to zero: same question in, same answer out. Grades are repeatable.மாதிரியின் சீரற்ற தன்மை பூஜ்ஜியமாக அமைக்கப்படுகிறது. ஒரே கேள்விக்கு ஒரே விடை வரும்; மதிப்பெண்களை மீண்டும் உருவாக்கலாம்.
letter-parseModels must answer A/B/C/D. Many write an essay first and bury the letter inside; the parser extracts the letter they committed to — or marks the answer wrong.மாதிரி A/B/C/D எழுத்தில் பதிலளிக்க வேண்டும். நீண்ட விளக்கம் எழுதினாலும், அதிலுள்ள இறுதி விடை எழுத்தை parser தேடும்; கிடைக்காவிட்டால் தவறு.
stratified sample · seed 42Every subject appears in proportion, and the fixed random seed means anyone re-running the bench gets the identical question set.ஒவ்வொரு பாடப்பிரிவும் உரிய விகிதத்தில் இடம்பெறும். seed 42 நிர்ணயிக்கப்பட்டதால் மீண்டும் இயக்குபவருக்கும் அதே கேள்வித் தொகுப்பு கிடைக்கும்.
nThe number of graded questions behind a score. Smaller n = wider confidence intervals.ஒரு மதிப்பெண்ணுக்குப் பின்னால் உள்ள மதிப்பிடப்பட்ட கேள்விகளின் எண்ணிக்கை. n குறைந்தால் நம்பிக்கை இடைவெளி பெரிதாகும்.

Every score on this site is tallied from saved answer sheets (results/*.jsonl). Three extracts, exactly as recorded:

இந்தத் தளத்தின் ஒவ்வொரு மதிப்பெண்ணும் சேமிக்கப்பட்ட விடைத்தாள்களிலிருந்து கணக்கிடப்படுகிறது. கீழே உள்ள மூன்று எடுத்துக்காட்டுகளும் பதிவான பதில்களே.

Paper 1 · MILU · Arts and Culture · gold: A · graded: correct (sample-milu-gemini.json — fresh call, exact bench prompt, recorded 2026-09-11)

Q: பின்வருபவர்களில் தாதா சாஹிப் பால்கே விருது 2013 பெற்ற நபர் யார்? (Who received the Dadasaheb Phalke Award for 2013?)
A. சம்பூரன் சிங் கல்ரா (Sampooran Singh Kalra — the poet Gulzar's real name) · B. விஜய ஷேசாஸ்திரி · C. ப்ரன் சிகந் · D. ரமேஷ் அகர்வால்
The model's entire reply: A — nothing else. The parser reads A, the gold is A, one mark. That is the whole contract: answer with a bare letter. When a model writes an essay instead, letter-parse digs the letter out of it — or marks the question wrong.

Paper 2 · IndicQA · question 1019 · Exact Match = 100, F1 = 100 (as recorded: results/indicqa_xiaomi_mimo-v2.5_n100.jsonl)

Passage: கோல்கொண்டா கோட்டை இந்திய தொல்லியல் துறை பட்டியலில் இடம் பெற்றுள்ளது… (The Golkonda fort features in the Archaeological Survey of India's list…)
Q: வெற்றியின் வாயில் எப்படி அழைக்கப்பட்டது? (What was the gate of victory called?) → Official answer: பதே தர்வாசா · MiMo v2.5 answered: பதே தர்வாசா — same words, wrapped in quotation marks. Normalization strips the punctuation, so it still scores word-for-word: full marks.

Paper 2 · IndicQA · EM = 0, F1 = 50 · the near-miss that explains both metrics (as recorded: results/indicqa_deepseek_deepseek-v4.1-flash_n100.jsonl)

Q: அன்னை தெரேசா எப்போது தனது சேவையைத் தொடங்கினார்? (When did Mother Teresa begin her service?) → Official answer: 1948 · DeepSeek V4.1 answered: 1948 ஆம் ஆண்டில் (“in the year 1948”). The model knows the answer — but adds a grammatical flourish, so exact-match fails while F1 pays half. Best part: all nine models flunked this exact question the exact same way. Multiply it across the paper and you have the whole EM-vs-F1 gap.

வினா 7

What the marks tell usமதிப்பெண்கள் சொல்லும் கதை

[5 × 2 = 10 marks]

இந்த முடிவுகள் நமக்குச் சொல்வது:

வினா 8

Run it yourselfநீங்களே இயக்கிப் பாருங்கள்

[practical]

The whole bench is one Python file, no GPU, no fine-tuning: it calls models through an OpenAI-compatible API exactly the way a normal user would. Set OPENROUTER_API_KEY, then:

இந்த benchmark ஒரு Python கோப்பு மட்டுமே; GPU அல்லது fine-tuning தேவையில்லை. சாதாரண பயனர் பயன்படுத்துவது போல OpenAI-compatible API வழியாக மாதிரிகளை அழைக்கிறது. OPENROUTER_API_KEY-ஐ அமைத்து, கீழே உள்ள கட்டளைகளை இயக்குங்கள்:

# Paper 1 — MILU MCQs (auto-downloads after you accept the HF gate)
python3 bench.py milu --model google/gemini-3.8-flash --n 199

# Paper 2 — IndicQA reading comprehension
python3 bench.py indicqa --model google/gemini-3.8-flash --n 100

# Re-score any saved answer sheet
python3 bench.py score --task milu --file results/milu_google_gemini-3.8-flash_n199.jsonl