The Exponential Rise of AI Intelligence
Understanding AI Progress for Urban Planners
Massive Multitask Language Understanding (MMLU) is like a comprehensive professional exam for AI models.
In just 5 years, AI went from random guessing to exceeding human experts.
| Year | Model | MMLU Score | What This Means |
|---|---|---|---|
| 2020 | Small models | ~25% | Random guessing (4 choices/question) |
| 2020 | GPT-3 | 43.9% | Barely better than chance |
| 2023 | GPT-4 | 86.4% | Near expert-level performance |
| 2024 | GPT-4o | 88.7% | Matching human experts |
| 2024 | Claude 3.5 | 88.3% | Matching human experts |
| 2025 | GPT-4.1 | 90.2% | Exceeding human experts |
| 2025 | GPT-5 | 91.4% | Significantly above experts |
| Human Experts | 89.8% | Baseline | |
Score jumped from 44% to 86%
That's like going from barely passing to top of the class
Improvement: +42 percentage points
Score jumped from 86% to 91%
Now exceeding human experts
Crossed human baseline: +5 points above experts
In the time it takes to complete a typical General Plan update (5-7 years), AI capabilities have gone from "barely useful" to "expert-level" across virtually all planning-related knowledge domains.
The question isn't whether AI will impact planning practiceβit already has. The question is: How do we adapt?
"For 50+ years of AI research, models scored near random chance on tasks like this. Then in just 3 years (2020-2023), we went from 44% to 86%. That's the exponential curve everyone talks about."
"As of 2025, the best AI models now score higher than human experts on broad knowledge tests. This doesn't mean AI is 'smarter' than humans, but it does mean the knowledge base available to AI-assisted planning is now broader than any single expert."
"Unlike previous waves of technology that automated routine tasks, these models demonstrate reasoning and knowledge synthesis at a professional level. This fundamentally changes what's possible in planning practice."