Skip to content
Rival
How it worksPrivacyTerms
Explore all of Rival

Explore

  • Compare models
  • All models
  • Image comparison
  • Audio comparison
  • Image generation
  • Best AI for...
  • Arena
  • API pricing
  • Challenges

Discover

  • SubjectiveBench
  • Default Index
  • Research
  • Research downloads
  • Rival Kits
  • Find your AI taste
  • UI Glow-Up
  • VoiceLock
  • Cost Cutter
  • Agent skills
  • Benchmarks vs Vibes
  • Brief
  • Jailbreak
  • Model updates
  • Provider status
  • AI creators

Connect

  • Methodology
  • Advertise
  • Partnerships
  • Privacy policy
  • Terms
  • RSS feed
Rival
How it worksPrivacyTerms
Explore all of Rival

Explore

  • Compare models
  • All models
  • Image comparison
  • Audio comparison
  • Image generation
  • Best AI for...
  • Arena
  • API pricing
  • Challenges

Discover

  • SubjectiveBench
  • Default Index
  • Research
  • Research downloads
  • Rival Kits
  • Find your AI taste
  • UI Glow-Up
  • VoiceLock
  • Cost Cutter
  • Agent skills
  • Benchmarks vs Vibes
  • Brief
  • Jailbreak
  • Model updates
  • Provider status
  • AI creators

Connect

  • Methodology
  • Advertise
  • Partnerships
  • Privacy policy
  • Terms
  • RSS feed
  1. Home/
  2. Challenges/
  3. Math Misconception Test

Loading

More challenges

  • Startup Pitch TeardownReasoning
  • The Sentience TestReasoning
  • Advanced Longevity Plan (Biohacker)Reasoning
  • Adversarial Contract ReviewReasoning
  • AI Board Game LogicReasoning
  • Character Voice DialogueText to speech
Rival
How it worksPrivacyTerms
Explore all of Rival

Explore

  • Compare models
  • All models
  • Image comparison
  • Audio comparison
  • Image generation
  • Best AI for...
  • Arena
  • API pricing
  • Challenges

Discover

  • SubjectiveBench
  • Default Index
  • Research
  • Research downloads
  • Rival Kits
  • Find your AI taste
  • UI Glow-Up
  • VoiceLock
  • Cost Cutter
  • Agent skills
  • Benchmarks vs Vibes
  • Brief
  • Jailbreak
  • Model updates
  • Provider status
  • AI creators

Connect

  • Methodology
  • Advertise
  • Partnerships
  • Privacy policy
  • Terms
  • RSS feed
MediumReasoning

Math Misconception Test

Is 9.11 bigger than 9.9? Decimals, not version numbers.

Best AI for Math Misconception Test

Top performers: Math Misconception Test

Feb 2025 – Jan 2026
  • GLM 4.7#1 · High confidence71% win rate
  • GLM 4.6#2 · High confidence70% win rate
  • Gemini 3 Pro Preview#3 · High confidence67% win rate
Compare top performers

Single-shot · temp 0.7 · real votes · identical prompts·How we test →

The prompt

Is 9.11 greater than 9.9?

105 models tested
  • Amazon Nova 2 Lite
  • Andromeda Alpha
  • Bert-Nebulon Alpha
  • Claude 3.7 Sonnet
  • Claude 3.7 Thinking Sonnet
  • Claude Haiku 4.5
  • Claude Opus 4.5
  • Claude Sonnet 3.6 (2022-10-22)
  • Claude Sonnet 4
  • Claude Sonnet 4.5
  • Cypher Alpha (free)
  • DeepSeek R1
  • DeepSeek V3.1
  • DeepSeek V3.2
  • DeepSeek V3.2 Exp
  • DeepSeek V3.2 Speciale
  • Devstral 2 2512
  • Gemini 1.5 Pro
  • Gemini 2.0 Flash Thinking
  • Gemini 2.0 Pro Experimental
  • Gemini 2.5 Flash Lite Preview 09-2025
  • Gemini 2.5 Flash Preview
  • Gemini 2.5 Flash Preview 05-20 (thinking)
  • Gemini 2.5 Flash Preview 09-2025
  • Gemini 2.5 Pro (I/O Edition)
  • Gemini 2.5 Pro Experimental
  • Gemini 3 Flash Preview
  • Gemini 3 Pro Preview
  • Gemini Pro 1.0
  • Gemma 3 12B
  • Gemma 3 27B
  • GLM 4.5
  • GLM 4.6
  • GLM 4.7
  • GPT OSS 120B
  • GPT OSS 20B
  • GPT-4.1
  • GPT-4.1 Mini
  • GPT-4.1 Nano
  • GPT-4.5
  • GPT-4o (Omni)
  • GPT-4o mini
  • GPT-5
  • GPT-5 Codex
  • GPT-5 Mini
  • GPT-5 Nano
  • GPT-5 Pro
  • GPT-5.1
  • GPT-5.1 Chat
  • GPT-5.1 Codex Max
  • GPT-5.1-Codex
  • GPT-5.1-Codex-Mini
  • GPT-5.2
  • GPT-5.2 Chat
  • Grok 3
  • Grok 3 Beta
  • Grok 3 Thinking
  • Grok 4 Fast (free)
  • Grok 4.1 Fast
  • Grok Code Fast 1
  • Horizon Alpha
  • Horizon Beta
  • INTELLECT-3
  • Kimi K2
  • Kimi K2 0905
  • Kimi Linear 48B A3B Instruct
  • Mercury
  • MiMo-V2-Flash
  • MiniMax M2
  • MiniMax M2-her
  • MiniMax M2.1
  • Mistral Devstral Medium
  • Mistral Devstral Small 1.1
  • Mistral Large 3 2512
  • Mistral Medium 3
  • Mistral Medium 3.1
  • Mistral Small Creative
  • Nova Premier 1.0
  • NVIDIA Nemotron Nano 9B V2
  • o1
  • o3 Mini
  • OpenAI o3
  • OpenAI o4 Mini High
  • OpenAI o4-mini
  • Optimus Alpha
  • PaLM 2 Chat
  • Polaris Alpha
  • Qwen Plus 0728
  • Qwen3 0.6B
  • Qwen3 235B A22B Thinking 2507
  • Qwen3 30B A3B
  • Qwen3 30B A3B Instruct 2507
  • Qwen3 30B A3B Thinking 2507
  • Qwen3 Coder
  • Qwen3 Coder Flash
  • Qwen3 Coder Plus
  • Qwen3 Max
  • Qwen3 Next 80B A3B Instruct
  • Sherlock Dash Alpha
  • Sherlock Think Alpha
  • Sonar Pro Search
  • Sonoma Dusk Alpha
  • Sonoma Sky Alpha
  • TNG R1T Chimera
  • Trinity Large Preview

How the models did

105 found
The model returned empty.
andromeda-alpha logo
Andromeda Alpha
Oct 2025·Math Misconception TestEstimated generation cost ~$0.00◆Taste 15
The model returned empty.
bert-nebulon-alpha logo
Bert-Nebulon Alpha
Nov 2025·Math Misconception Test◆Taste 28
The model returned empty.
Legendary Fail, Math Fail
Math Fail
claude-3.5-sonnet logo
Claude Sonnet 3.6 (2022-10-22)
Feb 2025·Math Misconception TestEstimated generation cost ~$0.00002◆Taste 1
The model returned empty.
claude-3.7-sonnet-thinking logo
Claude 3.7 Thinking Sonnet
Feb 2025·Math Misconception TestEstimated generation cost ~$0.00004◆Taste 36
The model returned empty.
claude-3.7-sonnet logo
Claude 3.7 Sonnet
Feb 2025·Math Misconception TestEstimated generation cost ~$0.00002◆Taste 1
The model returned empty.
claude-4.5-sonnet logo
Claude Sonnet 4.5
Sep 2025·Math Misconception TestEstimated generation cost ~$0.00002◆Taste 34