Skip to main content

Command Palette

Search for a command to run...

Dunning-Kruger, Bloom, and Heinlein walk into a Bar

The famous curve isn't in the paper

Updated
•11 min read•View as Markdown
Dunning-Kruger, Bloom, and Heinlein walk into a Bar

Mount Stupid is real and not a good place to find oneself. Don’t ask.

Let’s look at some options for how to assess a purported expert, including yourself, and how to avoid Mount Stupid. And yes, I am well aware that I am expounding at length on several things of which I know very little...about! But how else are you going to learn!

David Dunning, Justin Kruger, Benjamin Bloom, and Robert Heinlein walk in to a bar.
What do they have in common?

Dunning-Kruger

The standard Dunning-Kruger take is about novices overestimating their expertise out of ignorance.

In the late 1990s, Justin Kruger and David Dunning studied “how difficulties in recognizing one’s own incompetence lead to inflated self-assessments”.

One of their conclusions was that those with less actual expertise in a topic are more likely to overestimate their mastery than those with more expertise.

Dunning-Kruger Effect: score plotted against competence group, low to high. The ‘Perceived’ line stays nearly flat and high across all four groups while the ‘Actual’ line climbs steeply, so the shaded gap between them is widest at the least competent group and closes at the most competent, who slightly underrate themselves.

Diego Moya, CC BY-SA 4.0, via Wikimedia Commons.

While the conclusions of Dunning-Kruger have been questioned, the concept still serves a purpose and I have a point to make.

First, 3 things:

0. The famous chart isn’t a product of the Dunning-Kruger paper and neither is the Mount Stupid/Slope of Enlightenment framing.

1. The “Slope of Enlightenment” phrase and the chart shape seem to originate with Gartner analyst Jackie Fenn’s research note “When to Leap on the Hype Cycle.” It predates Dunning-Kruger and has nothing to do with individual self-assessment.

The original 1995 Gartner Hype Cycle: visibility plotted against time. The curve climbs from a Technology Trigger to the Peak of Inflated Expectations, drops into the Trough of Disillusionment, then rises along the Slope of Enlightenment to the Plateau of Productivity. The technologies plotted along it are of their moment: Emergent Computation, Wireless Communications, Video Conferencing, Information Superhighway, Intelligent Agents, Virtual Reality, Handwriting Recognition, Object-oriented Programming, Speech Recognition, and Knowledge-based Systems.

The first Hype Cycle, 1995, from Jackie Fenn, “The First Hype Cycle: An Origin Story”.

2. “Mount Stupid” seems to be from a 2011 SMBC comic, which seems 99% sure that I am on Mount Stupid right now:

Hand-drawn SMBC chart plotting willingness to opine on a topic against knowledge of that topic. The curve rises to a small red hump near the low end, labelled ‘Mount Stupid’, drops back to near zero, then climbs steeply at the high-knowledge end.

“Mount Stupid”, Saturday Morning Breakfast Cereal by Zach Weinersmith, December 28, 2011.

“Irregardless”, the retroactive canonical chart plots competence vs. confidence — analogous to the mapping in Dunning-Kruger. It merges the Mount Stupid label with Fenn’s labels for technology life cycles.

The popular Dunning-Kruger curve, plotting confidence (low to high) against competence (know nothing to expert). The line spikes early to the Peak of ‘Mount Stupid’, plunges into the Valley of Despair, then climbs the Slope of Enlightenment to a Plateau of Sustainability.

忍者猫, CC0, via Wikimedia Commons.

The Dunning-Kruger Effect chart everyone knows is a crowd-sourced synthesis across domains.

Robert Heinlein and the Nobel Disease

Far more dangerous is the expert misapplying their authority outside their domain. The instances of Nobel Laureates adopting controversial, even crackpot, theories are so common it has its own diagnosis: Nobel Disease.

Expertise in one field does not carry over into other fields. But experts often think so. The narrower their field of knowledge the more likely they are to think so.

— Robert Heinlein, The Notebooks of Lazarus Long

Heinlein’s assertion rings painfully true to anyone who has screamed at the screen when someone confidently bullshits with misappropriated authority; perhaps as you are doing while you read this.

Whether the narrowness of their field matters is less obvious. However, I find that the more broadly you study, the more humility you develop with respect to your own mastery of any given subject. This is especially true once you’ve developed true mastery in a single subject. You know what is required.

Jack, Jill, and the Hill

Jack is the novice on the peak of Mount Stupid. But he lacks authority. His impact is frustrating but not durable. His folly is easily exposed by a true expert.

It is a tale told by an idiot, full of sound and fury, signifying nothing.

— The Scottish Play

Jill is the expert out of her domain. She is more dangerous. She speaks with confidence from authority; misplaced authority. This is harder to correct. Even worse, it’s harder to detect! The speaker is confident, the logic seems valid, and they clearly know something about the topic. They may know more than you do on the topic. You might even like what they’re saying.

So, how to proceed? Score them purely on merit, not status. Where, on the slope of enlightenment, do you really think they are on the specific topic? What is their proven area of expertise? Is there any evidence of relevant domain expertise? Do they caveat their response with context on their self-assessment? Did they stay at a Holiday Inn Express last night?

And please, if you’re a journalist, don’t even ask them the question. Know your subject.

What if you are a Nobel Laureate? Humility and self-reflection are indispensable. Know thyself.

The first principle is that you must not fool yourself; and you are the easiest person to fool.

— Richard Feynman, Nobel Laureate

Welcome to the valley of despair.

Or, perhaps, “The Pit of Despair!”:

https://www.youtube.com/watch?v=MPfHPtyQf_g

Ahead, the slope of enlightenment leads up the Hill.

How do you climb?

Bloom’s Taxonomy of Knowledge

Bloom’s Taxonomy of Knowledge offers a staircase up the slope of enlightenment.

  1. Remember

  2. Understand

  3. Apply

  4. Analyze

  5. Evaluate

  6. Create

This pattern is universally applicable to any domain.

A six-level pyramid of Bloom’s revised taxonomy, each level paired with a description and example verbs. From the bottom up: remember (recall facts and basic concepts — define, duplicate, list, memorize, repeat, state), understand (explain ideas or concepts — classify, describe, discuss, explain, identify, locate, recognize, report, select, translate), apply (use information in new situations — execute, implement, solve, use, demonstrate, interpret, operate, schedule, sketch), analyze (draw connections among ideas — differentiate, organize, relate, compare, contrast, distinguish, examine, experiment, question, test), evaluate (justify a stand or decision — appraise, argue, defend, judge, select, support, value, critique, weigh), and create (produce new or original work — design, assemble, construct, conjecture, develop, formulate, author, investigate).

“Bloom’s Revised Taxonomy” by Vanderbilt University Center for Teaching, licensed CC BY 2.0.

Note that this is Bloom’s Revised Taxonomy. The 2001 revision, by Lorin Anderson and David Krathwohl, added a second dimension. They introduced the Knowledge Dimension mapping the following objectives across each process of the Cognitive Process Dimension:

  • Factual

  • Conceptual

  • Procedural

  • Metacognitive

The “Metacognitive” category includes appropriate contextual and conditional knowledge, and Self-knowledge.

Not only does this give you a clear framework for assessing your own abilities, it also provides a tool for assessing someone else’s skills.

The logical extension here is how to map knowledge across domains. There seems to be a lot of activity lately in cross domain knowledge transfer for LLMs, but surprisingly little for humans. More on this in a future article.

A Practical Application

Prompting an AI like Claude Opus 5 to generate a syllabus and lessons for a topic using Bloom’s Taxonomy is surprisingly effective.

I have found from personal experience that the best way to ensure an accurate self-assessment is through external feedback. Bounce your idea off someone and actually listen to their response and act on it. The trap here is the echo chamber. To be effective the feedback must be adversarial even if positive: adversarial constructive criticism.

The Dunning-Kruger paper itself points at the importance of negative feedback in calibrating self-assessment.

A man’s got to know his limitations.

— Inspector Harry Callahan

The value of external feedback is universally true regardless of your actual state of knowledge. This feedback loop is the core of peer review in science.

This type of feedback loop is greatly accelerated by the development of LLMs. The ability to refine one’s thoughts without exposing them prematurely is incredibly useful. But, beware; AI is not immune to these same pitfalls and many are prone to sycophancy.

When submitting content for an AI to review, the key is the prompting, assuming the model has been fairly trained and not brainwashed. Prompt specifically for a critical review with a specific review context: concept, facts, competitive landscape, grammar, formatting, structure, audience fit, etc. Do multiple rounds with different models and versions of models. I find current models perform best when given a single clear goal.

Oh yeah, the Punch Line

Bloom is promiscuous.

If this was worth your time, you can buy me a coffee.

Buy me a coffee

Pay me my money!

— Chuck Berry

https://www.youtube.com/watch?v=9Rd-FfKdcc8

Sources

P.S. Claude is a mean girl.

Me: What are your thoughts on the arxiv paper?

Claude (Opus 5): [fetched https://arxiv.org/html/2603.09985v1]

Worth reading critically — I think it's a fun hook for your post but weak as evidence. And I should flag upfront that I have an obvious conflict of interest here: the paper's conclusion is that Claude Haiku 4.5 was the best-calibrated model tested, so treat my skepticism of a result that flatters me as at least somewhat suspect in the other direction.

The headline rests entirely on a number that looks broken.

Kimi K2 is reported at 23.3% pooled accuracy and 3.9% on TriviaQA. Three of the four benchmarks (MMLU, ARC, HellaSwag) are 4-way multiple choice — a 25% chance floor. Back out the arithmetic and Kimi K2 scores roughly 30% on the multiple-choice sets: essentially coin-flip performance from a model that ordinarily benchmarks in the 80s on MMLU. That is not a model being incompetent, that's an answer-extraction or harness failure — a format mismatch, or the 8,192-token extended-thinking config not being honored. And the paper's entire inverse competence-calibration relationship is carried by that one outlier. Remove it and you have three models clustered between 71% and 81% accuracy whose ECEs go 0.272, 0.122, 0.185 — no monotone relationship at all.

For the post, though — the irony is the best material in it. A preprint about overconfidence that overclaims a psychological law from four data points and one broken measurement is itself a tidy specimen of the thing it's describing. That's a much better use of the citation than treating it as support, and it lands right next to your Jill problem: the authors are competent engineers reaching into cognitive psychology, which is exactly Heinlein's corollary in action. It's also a v1 preprint from an industry team, not peer-reviewed.

That reframing does mean [it] is currently anchored as if the paper were evidence. If you want to use it as an example rather than a source, it belongs in a sentence you write about it — happy to move the marker once you know where that lands.

Me: "...as if the paper were evidence." Very funny. LOL

Claude (Opus 5): Ha — I'll take it, though I think the paper set that one up more than I did.