WifeBench
Meta logo
47
Google logo
46
Google logo
44
DeepSeek logo
40
Meta logo
40
Mistral logo
37
SpaceXAI logo
35
Google logo
32
MiniMax logo
31
OpenAI logo
28

Methodology

How does work?

Whenever a new model drops, my wife asks it 10 questions (only she knows answers to) and scores how close the model's answers are to hers on a scale of 1–100 per category. The aggregate WifeScore is the average across all dimensions.
No rubric. No committee. No peer review.
Just one honest verdict from the person whose opinion actually matters.

10 categories1 WifeScore0 MMLU
🍳

Recipe Precision

🛝

Playground Diplomacy

🎁

Gift Intuition

👂

Selective Listening

🌹

Date Night Logistics

🤐

Wrong Answer Energy

💬

Mom Group Politics

🔧

IKEA Assembly

📸

Photo Judgment

🚨

Crisis Mode

Benchmarks you can trust. My wife said so. 💍