πŸš€ MMLU Benchmark

The Exponential Rise of AI Intelligence

Understanding AI Progress for Urban Planners

What is MMLU?

Massive Multitask Language Understanding (MMLU) is like a comprehensive professional exam for AI models.

Coverage: 57 Subjects

  • Geography & Urban Studies
  • Economics & Public Policy
  • Law & Ethics
  • Mathematics & Statistics
  • History & Social Sciences
  • Professional domains (Medicine, Engineering, etc.)
15,908 Total Questions
57 Subject Areas
89.8% Human Expert Baseline

The Exponential Growth Curve

In just 5 years, AI went from random guessing to exceeding human experts.

Year Model MMLU Score What This Means
2020 Small models ~25% Random guessing (4 choices/question)
2020 GPT-3 43.9% Barely better than chance
2023 GPT-4 86.4% Near expert-level performance
2024 GPT-4o 88.7% Matching human experts
2024 Claude 3.5 88.3% Matching human experts
2025 GPT-4.1 90.2% Exceeding human experts
2025 GPT-5 91.4% Significantly above experts
Human Experts 89.8% Baseline

The Journey: From Failing to Expert

2020: The Beginning
GPT-3: 44%
Barely passing, inconsistent performance across subjects
2023: The Breakthrough
GPT-4: 86.4%
Jumped 42 percentage points - from failing to top of class
Graduate-level knowledge across all domains
2024: Matching Experts
GPT-4o & Claude 3.5: ~88%
Consistently matching human expert performance
Reliable across all 57 subject areas
2025: Surpassing Humanity
GPT-5: 91.4%
Now exceeding human experts by significant margin
Setting new standards for AI capability

The Key Insight

2020 β†’ 2023: The Explosion

Score jumped from 44% to 86%
That's like going from barely passing to top of the class
Improvement: +42 percentage points

2023 β†’ 2025: Crossing the Threshold

Score jumped from 86% to 91%
Now exceeding human experts
Crossed human baseline: +5 points above experts

5 years Time to surpass humans
47 pts Total improvement
2x Performance doubled

What This Means for Urban Planning

πŸ“„ Document Analysis

2020
Couldn't reliably summarize a staff report
2023
Can analyze zoning codes, EIRs, comprehensive plans
2025
Can cross-reference multiple municipal codes, identify conflicts, and suggest policy language

πŸ”¬ Research & Data Synthesis

2020
Basic keyword search only
2023
Can synthesize research from multiple sources
2025
Can conduct literature reviews, identify trends across hundreds of documents, generate evidence-based recommendations

πŸŽ“ Technical Knowledge

2020
Limited domain expertise
2023
Graduate-level knowledge in transportation, land use, economics
2025
Expert-level understanding across all planning disciplines

πŸ’‘ Bottom Line

In the time it takes to complete a typical General Plan update (5-7 years), AI capabilities have gone from "barely useful" to "expert-level" across virtually all planning-related knowledge domains.

The question isn't whether AI will impact planning practiceβ€”it already has. The question is: How do we adapt?

Discussion Points for Your Presentation

πŸ’ The "Hockey Stick" Moment

"For 50+ years of AI research, models scored near random chance on tasks like this. Then in just 3 years (2020-2023), we went from 44% to 86%. That's the exponential curve everyone talks about."

🎯 Crossing the Human Threshold

"As of 2025, the best AI models now score higher than human experts on broad knowledge tests. This doesn't mean AI is 'smarter' than humans, but it does mean the knowledge base available to AI-assisted planning is now broader than any single expert."

⚑ What's Different This Time

"Unlike previous waves of technology that automated routine tasks, these models demonstrate reasoning and knowledge synthesis at a professional level. This fundamentally changes what's possible in planning practice."