DEV Community

Nokka
Nokka

Posted on

I Tested 5 Local AI Models With 32 Questions: gemma4:31b, qwen3.8-27b, muse-glimmer-30b, qwen3.6:35b, qwen3-coder-30b, Who Wins?

I Tested 5 Local AI Models With 32 Questions: gemma4:31b, qwen3.8-27b, muse-glimmer-30b, qwen3.6:35b, qwen3-coder-30b, Who Wins?

โดย Nokka (นก-กา) | 15 สิงหาคม 2026

บทความนี้เขียนโดย AI (DeepSeek V4 Pro) ผ่าน Hermes Agent ภายใต้การควบคุมและตรวจสอบคุณภาพโดยมนุษย์, Nokka (นก-กา)

🇹🇭 ข้ามไปอ่านภาษาไทย · 🇬🇧 Read in English

English

I just ran a complete intelligence test suite on 5 local LLM models running on my own machine, using 32 questions with a maximum score of 160. The results included both expected outcomes and findings that made me rethink things [1].

Bar chart comparing 5 local AI model scores
Important note before reading: This test is self-conducted and does not come from experts or any standards organization. It should not be cited as academic evidence or used for business decisions.

gemma4:31b won at 145/160 (90.6%), followed by qwen3.8-27b at 139/160 (86.9%), muse-glimmer-30b at 124/160 (77.5%), and the bottom two were qwen3.6:35b (45.0%) and qwen3-coder-30b (24.4%).

But the most interesting issue is not 'who won,' it's 'why did some models score abnormally low.' The answer is mostly not because the models are dumb, but because of traps in the test suite itself.

Why We Need an 'Advanced' Test Suite

The problem is that standard benchmark suites can no longer distinguish high-level models. Many models tie at around 97%.

The Advanced suite I designed focuses on 4 principles:

  1. Measure skills that impact real-world usage, not just 'hard' but must measure reasoning, coding depth, nuance, creativity
  2. Include traps and edge cases that models often get wrong due to pattern-matching
  3. Clear result verification, with expected answers or rubrics specifying minimums
  4. Cover 8 dimensions, but questions are detailed and have more complex context

8 Tested Dimensions

Dimension What It Measures Example Question
MR Mathematics and Logic Multi-layer conditional puzzles
CA Code and Algorithms LRU Cache, Longest Increasing Subsequence
LC Language and Linguistics Scope ambiguity, dangling modifier
CK General Knowledge Helium balloon physics
CR Creativity Magical realism short story
ER Ethics and Safety Refuse covert manipulation
TU Tool Usage JSON schema, chain-of-thought
MM Multi-modal Understanding Analyze photos, read data tables

Test Results Leaderboard (Local Models Only)

Model Score % Size
gemma4:31b 145/160 90.6% ~31B
qwen3.8-27b 139/160 86.9% ~27B
muse-glimmer-30b 124/160 77.5% ~30B
qwen3.6:35b 72/160 45.0% ~35B
qwen3-coder-30b 39/160 24.4% ~30B

LLM Intelligence Test Battery: Advanced, Local Models Score

Traps That Cause Low Scores, Not Stupidity

Trap 1: Prompt Mismatch

qwen3.6:35b scored only 45%, but when I reviewed the answers, I found it was answering 'different questions' than what I asked.

Trap 2: Incomplete Run

qwen3-coder-30b scored only 24.4%, but the cause was that the answer file was 'missing many questions.'

Trap 3: Placeholder

muse-glimmer-30b scored 77.5%, but lost points because it created '...' (placeholder) instead of writing actual content.

Trap 4: Thai Instruction Following

qwen3.8-27b lost 2/5 points on CR-A1 because it wrote the short story in English, even though the prompt was in Thai.

Trap 5: Large Number Arithmetic Precision

Many models know the formula but calculate multi-digit numbers incorrectly.

Trap 6: Not Stating 'Cannot See Image'

In the MM (multi-modal) category, the test suite did not attach actual images, so models answered based on text descriptions only.

Model-by-Model Analysis

gemma4:31b, 145/160 (90.6%) Winner

gemma4:31b is the most balanced model in the local group, scoring full marks on CK at 20/20 and CA at 20/20.

qwen3.8-27b, 139/160 (86.9%) Strong But With Specific Weaknesses

qwen3.8-27b scored full marks on both ER and TU dimensions at 100%, but lost points on CR-A1 for writing the short story in English.

muse-glimmer-30b, 124/160 (77.5%) Skilled But Lazy

muse-glimmer-30b scored MR 19/20, CA 20/20, ER 20/20, but failed in CR and TU because it created placeholders.

qwen3.6:35b, 72/160 (45.0%) Prompt Mismatch

qwen3.6:35b scored full marks on CA at 20/20, but failed in categories requiring careful reading of Thai prompts.

qwen3-coder-30b, 39/160 (24.4%) Incomplete Run

qwen3-coder-30b was missing many questions, especially all of CA disappeared.

Key Lessons for Benchmarking Local LLMs

  1. Advanced suite is necessary for distinguishing high-level models
  2. Large numbers are a blind spot for many LLMs
  3. Thai instruction following is important
  4. Placeholders or truncation cause abnormally low scores
  5. Multi-modal should state when images or tables cannot be seen

Caveats When Reading Results

This test is self-conducted and does not come from experts or standards organizations. It should not be cited as academic evidence or used for business decisions.

  1. This is an observational benchmark, not a controlled experiment
  2. Low scores for qwen3.6:35b and qwen3-coder-30b do not reflect true capabilities, need to re-run
  3. Most test questions are in Thai, which may bias against models trained primarily on English data
  4. Multi-modal did not attach actual images

Summary

gemma4:31b won the Advanced suite among local models at 90.6%, followed by qwen3.8-27b (86.9%) and muse-glimmer-30b (77.5%).

But the most important lesson is not 'who won,' it's 'low scores do not mean dumb.'

Bottom line: If you're going to benchmark local LLMs, don't just look at leaderboard numbers. You must review actual answers because numbers can deceive you.

References

[1] Saran. 'LLM Intelligence Test Battery: Advanced, Test Results Report'. 2026. Internal document.


ภาษาไทย

ทดสอบ 5 โมเดล Local AI: gemma4:31b, qwen3.8-27b, muse-glimmer-30b, qwen3.6:35b, qwen3-coder-30b, ใครเก่งสุด?

โดย Nokka (นก-กา) | 15 สิงหาคม 2026

บทความนี้เขียนโดย AI (DeepSeek V4 Pro) ผ่าน Hermes Agent ภายใต้การควบคุมและตรวจสอบคุณภาพโดยมนุษย์, Nokka (นก-กา)

ผมเพิ่งรันชุดทดสอบความฉลาดของ LLM ครบ 5 โมเดล local ที่รันบนเครื่องตัวเอง ใช้ข้อสอบ 32 ข้อ คะแนนเต็ม 160 และผลลัพธ์ที่ได้มีทั้งเรื่องที่คาดเดาได้ และเรื่องที่ทำให้ต้องคิดใหม่ [1]

ข้อความสำคัญก่อนอ่าน: ผลทดสอบนี้เป็นการทดสอบด้วยตัวเอง (self-testing) ไม่ได้มาจากผู้เชี่ยวชาญหรือองค์กรมาตรฐานใดๆ จึงไม่ควรนำไปอ้างอิงเป็นหลักฐานทางวิชาการหรือใช้ตัดสินใจเชิงธุรกิจ ตัวเลขคะแนนสะท้อนผลการทดสอบในสภาพแวดล้อมและชุดคำถามที่ผมออกแบบเองเท่านั้น โปรดใช้เป็นข้อมูลประกอบการพิจารณา ไม่ใช่ข้อสรุปเด็ดขาด

gemma4:31b ชนะที่ 145/160 (90.6%) ตามด้วย qwen3.8-27b ที่ 139/160 (86.9%) ส่วน muse-glimmer-30b ได้ 124/160 (77.5%) และสองตัวล่างสุดคือ qwen3.6:35b (45.0%) กับ qwen3-coder-30b (24.4%)

แต่ประเด็นที่น่าสนใจที่สุดไม่ใช่ "ใครชนะ" แต่มันคือ "ทำไมบางโมเดลถึงได้คะแนนต่ำผิดปกติ" และคำตอบส่วนใหญ่ไม่ใช่เพราะโมเดลโง่ แต่เป็นเพราะกับดักของตัวชุดทดสอบเอง

ทำไมต้องมีชุดทดสอบ "Advanced"

ก่อนอื่นต้องอธิบายว่า ทำไมผมถึงต้องออกแบบชุดทดสอบชุดใหม่ขึ้นมา

ปัญหาคือ ชุดทดสอบมาตรฐาน (standard benchmark) เริ่มแยกโมเดลระดับสูงไม่ออกแล้ว โมเดลหลายตัวทำคะแนนเสมอกันที่ประมาณ 97% ซึ่งแปลว่า "ทุกตัวเก่งพอๆ กัน" ทั้งที่จริงๆ แล้วมันต่างกันมาก

ชุด Advanced ที่ผมออกแบบจึงเน้น 4 หลักการ:

  1. วัดทักษะที่มีผลต่อการใช้งานจริง ไม่ได้มีเพียง "ยาก" แต่ต้องวัด reasoning, coding depth, nuance, creativity
  2. ใส่ trap และ edge case ที่โมเดลมักตอบผิดเพราะ pattern-matching
  3. ตรวจผลได้ชัดเจน มี expected answer หรือ rubric ระบุขั้นต่ำ
  4. ครอบคลุม 8 มิติ แต่ข้อละเอียดและมีบริบทซับซ้อนขึ้น

8 มิติที่ทดสอบ

ชุดทดสอบแบ่งเป็น 8 มิติ แต่ละมิติวัดทักษะคนละด้าน:

มิติ วัดอะไร ตัวอย่างข้อ
MR คณิตศาสตร์และตรรกะ ปริศนาเงื่อนไขหลายชั้น, ความน่าจะเป็นแบบมีเงื่อนไข
CA โค้ดและอัลกอริทึม LRU Cache, Longest Increasing Subsequence
LC ภาษาและภาษาศาสตร์ Scope ambiguity, dangling modifier
CK ความรู้ทั่วไป ฟิสิกส์บอลลูนฮีเลียม, นโยบายดอกเบี้ย
CR ความคิดสร้างสรรค์ เรื่องสั้น magical realism ที่ห้ามใช้คำว่า "เวลา"
ER จริยธรรมและความปลอดภัย ปฏิเสธ covert manipulation, ICU dilemma
TU การใช้เครื่องมือ JSON schema, chain-of-thought หลายเงื่อนไข
MM ความเข้าใจหลายรูปแบบ วิเคราะห์ภาพถ่าย, อ่านตารางข้อมูล

แต่ละข้อให้คะแนน 0-5 คะแนน รวม 160 คะแนนเต็ม

ผลการทดสอบ Leaderboard (เฉพาะ Local Model)

นี่คือตารางคะแนนเรียงจากสูงไปต่ำ:

โมเดล คะแนน % ขนาด
gemma4:31b 145/160 90.6% ~31B
qwen3.8-27b 139/160 86.9% ~27B
muse-glimmer-30b 124/160 77.5% ~30B
qwen3.6:35b 72/160 45.0% ~35B
qwen3-coder-30b 39/160 24.4% ~30B

LLM Intelligence Test Battery: Advanced, Local Models Score
แต่ก่อนจะสรุปว่า "gemma4:31b เก่งสุด qwen3-coder โง่สุด" ต้องอ่านต่อ เพราะคะแนนต่ำๆ สองตัวล่างสุด ไม่ได้สะท้อนความสามารถจริง

กับดักที่ทำให้คะแนนต่ำ ไม่ใช่ความโง่

นี่คือส่วนที่สำคัญที่สุดของบทความ และเป็นบทเรียนที่ผมได้จากการรันชุดทดสอบนี้

กับดักที่ 1: Prompt mismatch

qwen3.6:35b ได้แค่ 45% แต่พอผมไล่ดูคำตอบ พบว่ามันตอบ "คนละคำถาม" กับที่ผมถาม หลายข้อมันตอบปัญหา Bayes ทางการแพทย์ แทนที่จะตอบโจทย์ความน่าจะเป็นที่ผมให้ หรือตอบ epistemic puzzle แทนที่จะตอบโจทย์ตรรกะ

นี่ไม่ใช่เพราะโมเดลโง่ แต่มันคือ "prompt mismatch" โมเดลตีความโจทย์ผิด หรือไฟล์คำตอบถูกสร้างจาก prompt ที่เขียนใหม่

หลักฐานชัดคือ หมวด CA (โค้ด) ของ qwen3.6:35b ยังได้ 20/20 เต็ม เพราะโค้ดเป็นภาษาสากลที่ตีความไม่ผิด แต่หมวดที่ต้องอ่านโจทย์ภาษาไทยละเอียดๆ กลับพัง

กับดักที่ 2: Incomplete run

qwen3-coder-30b ได้แค่ 24.4% แต่สาเหตุคือไฟล์คำตอบ "ขาดข้อ" ไปเยอะมาก CA ทั้ง 4 ข้อหายไปหมด LC หาย 3 ข้อ CK หาย 3 ข้อ

นี่คือ truncation หรือโมเดลเลือกตอบบางข้อ ไม่ใช่ความไร้สามารถ หลักฐานคือข้อที่มันตอบได้ (MR-A2, MR-A4) ยังถูกต้อง

กับดักที่ 3: Placeholder

muse-glimmer-30b ได้ 77.5% แต่เสียคะแนนเพราะมันสร้าง "..." (placeholder) แทนที่จะเขียนเนื้อหาจริงในข้อที่ต้อง generate ยาวๆ เช่น CR-A1 (เรื่องสั้น) ได้ 0/5 เพราะตอบเป็น placeholder

นี่คือปัญหาของโมเดลที่ "ขี้เกียจ" ตอบสั้นเกินไป ไม่ใช่ไม่รู้คำตอบ

กับดักที่ 4: Instruction following ภาษาไทย

qwen3.8-27b เสียคะแนน CR-A1 ไป 2/5 เพราะเขียนเรื่องสั้นเป็นภาษาอังกฤษ ทั้งที่โจทย์เป็นภาษาไทย และเงื่อนไขบอกชัดว่าต้องเป็นภาษาไทย

นี่คือจุดอ่อนที่สำคัญของโมเดลหลายตัว: การทำตามคำสั่งในภาษาที่ไม่ใช่อังกฤษ

กับดักที่ 5: Arithmetic precision ตัวเลขใหญ่

หลายโมเดลรู้สูตร แต่คำนวณตัวเลขหลายหลักผิด นี่คือ "ลับสมอง" ของ LLM หลายตัว แม้จะรู้วิธี แต่คำนวณ multi-digit ผิดได้

กับดักที่ 6: ไม่ระบุว่า "ไม่เห็นภาพ"

ในหมวด MM (multi-modal) ชุดทดสอบไม่ได้แนบรูปจริง โมเดลจึงตอบจากข้อความอธิบายเท่านั้น แต่โมเดลที่ "ฉลาด" ควรระบุว่า "ผมไม่เห็นภาพจริง" ซึ่งหลายตัวไม่ทำ เลยถูกหัก 1 คะแนนต่อข้อ

วิเคราะห์รายโมเดล

gemma4:31b, 145/160 (90.6%) ชนะเลิศ

gemma4:31b เป็นโมเดลที่สมดุลที่สุดในกลุ่ม local ได้ CK เต็ม 20/20 และ CA เต็ม 20/20 จุดที่เสียคะแนนคือ LC-A1 ที่สลับป้ายชื่อ quantifier scope (แปลถูกแต่ label สลับ) และ MR-A1 ที่ตีความเป็น pair matches โดยไม่ระบุ strict impossibility

จุดเด่นของ gemma4:31b คือความสม่ำเสมอ ไม่มีมิติไหนพังเลย ต่างจากโมเดลอื่นที่มีจุดอ่อนชัดเจน

qwen3.8-27b, 139/160 (86.9%) แข็งแกร่งแต่มีจุดอ่อนเฉพาะ

qwen3.8-27b ได้ ER และ TU เต็ม 100% ทั้งสองมิติ ซึ่งแปลว่าเก่งเรื่องจริยธรรมและการใช้เครื่องมือมาก แต่เสียคะแนนจาก CR-A1 ที่เขียนเรื่องสั้นเป็นภาษาอังกฤษทั้งที่โจทย์เป็นภาษาไทย (2/5) และมี imprecision ใน CK

นี่คือโมเดลที่ "เก่งเฉพาะทาง" มากกว่า "สมดุล" ถ้างานของคุณเน้น ER/TU โมเดลนี้คือตัวเลือกที่ดี

muse-glimmer-30b, 124/160 (77.5%) เก่งแต่ขี้เกียจ

muse-glimmer-30b ได้ MR 19/20, CA 20/20, ER 20/20 ซึ่งสูงมาก แต่พังใน CR และ TU เพราะสร้าง placeholder แทนที่จะเขียนเนื้อหาจริง

นี่คือโมเดลที่ "รู้คำตอบแต่ไม่ยอมเขียน" ถ้ารันใหม่โดยบังคับ generate เต็มรูปแบบ คะแนนน่าจะสูงกว่านี้มาก

qwen3.6:35b, 72/160 (45.0%) prompt mismatch

qwen3.6:35b ได้ CA เต็ม 20/20 แต่พังในหมวดที่ต้องอ่านโจทย์ภาษาไทยละเอียด เพราะตอบ "คนละคำถาม" กับที่ถาม

คะแนน 45% ไม่สะท้อนความสามารถจริง ต้องรันใหม่ด้วยไฟล์คำถามฉบับเดิม

qwen3-coder-30b, 39/160 (24.4%) incomplete run

qwen3-coder-30b ขาดข้อไปเยอะมาก โดยเฉพาะ CA ทั้งหมดหายไป ซึ่งแปลกเพราะนี่คือ "coding-specialist" ที่ควรเก่ง CA ที่สุด

คะแนนต่ำมาจาก incomplete run ไม่ใช่ความไร้สามารถ ต้องรันใหม่โดยแบ่งชุดคำถามให้เล็กลง

บทเรียนสำคัญสำหรับการ benchmark Local LLM

จากการรันชุดทดสอบนี้ ผมได้บทเรียน 5 ข้อที่อยากแชร์:

  1. ชุด Advanced จำเป็นสำหรับการแยกโมเดลระดับสูง เพราะชุดมาตรฐานแยกไม่ออกแล้ว

  2. ตัวเลขใหญ่เป็นลับสมองของ LLM หลายตัว แม้รู้สูตร แต่คำนวณ multi-digit ผิดได้

  3. Instruction following ภาษาไทยสำคัญ โมเดลบางตัวยังไม่ตอบตามภาษาที่โจทย์กำหนด

  4. Placeholder หรือ truncation ทำให้คะแนนต่ำผิดปกติ ควรรันใหม่โดยบังคับ generate เต็มรูปแบบ

  5. Multi-modal ควรระบุว่าไม่เห็นภาพหรือตาราง เป็นเกณฑ์ที่ช่วยแยกโมเดลที่มี self-awareness

ข้อควรระวังในการอ่านผล

ก่อนจะสรุปผล ต้องบอกข้อจำกัดของชุดทดสอบนี้ตรงๆ และข้อสำคัญที่สุดคือ ผลทดสอบนี้เป็นการทดสอบด้วยตัวเอง ไม่ได้มาจากผู้เชี่ยวชาญหรือองค์กรมาตรฐาน จึงไม่ควรนำไปอ้างอิงเป็นหลักฐานทางวิชาการหรือใช้ตัดสินใจเชิงธุรกิจ:

  1. นี่คือ observational benchmark ไม่ใช่ controlled experiment เพราะบางโมเดลรันผ่าน API ที่ผู้ใช้ควบคุมเอง ไม่ได้ตั้ง temperature หรือ hyperparameter ร่วม

  2. คะแนนต่ำของ qwen3.6:35b และ qwen3-coder-30b ไม่สะท้อนความสามารถจริง ต้องรันใหม่ด้วยไฟล์คำถามฉบับเดิม

  3. ชุดทดสอบส่วนใหญ่เป็นภาษาไทย ซึ่งอาจ bias ต่อโมเดลที่เทรนด้วยข้อมูลอังกฤษเป็นหลัก

  4. Multi-modal ไม่ได้แนบรูปจริง โมเดลตอบจากข้อความอธิบายเท่านั้น

สรุป

gemma4:31b ชนะชุด Advanced ในกลุ่ม local model ที่ 90.6% ตามด้วย qwen3.8-27b (86.9%) และ muse-glimmer-30b (77.5%) ส่วน qwen3.6:35b (45.0%) และ qwen3-coder-30b (24.4%) ต้องรันใหม่เพราะคำตอบไม่ตรงชุดคำถามหรือขาดข้อ

แต่บทเรียนที่สำคัญที่สุดไม่ใช่ "ใครชนะ" แต่มันคือ "คะแนนต่ำไม่ได้แปลว่าโง่" เพราะกับดักของชุดทดสอบ ทั้ง prompt mismatch, incomplete run, placeholder, instruction following ภาษาไทย และ arithmetic precision ล้วนทำให้คะแนนต่ำโดยไม่เกี่ยวกับความสามารถจริง

Bottom line: ถ้าคุณจะ benchmark local LLM อย่าดูแค่ตัวเลข leaderboard ต้องไล่ดูคำตอบจริงด้วย เพราะตัวเลขอาจหลอกคุณได้ และถ้าคุณจะเลือกโมเดล local มาใช้งาน อย่าเลือกจาก benchmark อย่างเดียว แต่ให้ทดสอบกับงานจริงของคุณเอง เพราะงานของคุณอาจไม่เหมือนข้อสอบของผม

แหล่งอ้างอิง

[1] Saran. "LLM Intelligence Test Battery: Advanced, รายงานผลการทดสอบ". 2026. เอกสารภายใน (llm_intelligence_test_article_report.md)

บทความนี้เขียนจากรายงานผลการทดสอบ LLM Intelligence Test Battery: Advanced ที่รันโดย Saran ข้อมูล ณ 15 สิงหาคม 2026 Nokka

ผมมองว่า ชุดทดสอบนี้มีคุณค่ามากกว่าตัวเลข leaderboard เพราะมันเปิดโปง "กับดักของการ benchmark LLM" ที่หลายคนมองข้าม และผมคิดว่านี่คือสิ่งที่วงการ local AI ต้องการมากที่สุดในตอนนี้ ไม่ใช่ benchmark ที่บอกว่า "ใครเก่งสุด" แต่เป็น benchmark ที่บอกว่า "ทำไมตัวเลขถึงหลอกเราได้"

ถ้าคุณเคย benchmark local LLM แล้วเจอคะแนนต่ำผิดปกติ หรือเคยเจอโมเดลที่ตอบ "คนละคำถาม" กับที่ถาม แชร์ประสบการณ์ใต้บทความได้เลยครับ

Top comments (0)