Calibration is the meeting where your ratings meet everyone else's. You walk in with "exceeds" for the engineer who carried the migration, someone across the table asks "compared to what?", and you find yourself telling a story about how hard they worked. Ten minutes later the rating is "meets", and you have to explain that to a person who trusted you to represent them. Most ratings that get moved in calibration don't get moved because they were wrong. They get moved because the manager couldn't show why they were right. This post covers how to prepare a case that survives the room, how to check your own ratings first, and how to behave once you're in there.
What calibration is for
Calibration exists because managers rate differently. Some are generous, some are strict, and most rate people who resemble them a little higher than people who don't. The meeting is meant to apply one standard across teams, so that "exceeds" means the same thing whichever manager you happen to report to.
That problem is bigger than most managers assume. Steven Scullen, Michael Mount and Maynard Goff studied multi-rater performance ratings of several thousand managers (published in the Journal of Applied Psychology in 2000) and found that the rater's own idiosyncratic tendencies explained more of the variation in ratings than the performance of the person being rated. Daniel Kahneman, Olivier Sibony and Cass Sunstein devote a chapter of Noise (2021) to performance ratings and come to a similar conclusion: much of what a rating measures is the rater. One of their remedies is to make the scale concrete, anchoring ratings to shared descriptions and examples rather than to adjectives like "strong" that each manager reads differently.
So treat calibration as a check on your judgment, not an attack on your team. If you walk in assuming your ratings are accurate and everyone else's are inflated, you'll argue badly and learn nothing.
What calibration should not be is forced ranking with a quota of low ratings. If your company runs a hard distribution, you still calibrate against the level expectations first; the distribution conversation comes after, and you should know which of the two you're having.
Build the case before the meeting
Each person you're rating needs a short written case. One page at most; half a page is better. Write it before you see anyone else's ratings, so you're anchored on evidence rather than on what other managers proposed.
A template:
Name, level, proposed rating
Expectations for this level (one line, quoted from your career ladder)
Top three contributions this period
For each: what they did, what changed because of it, and how you know.
Name the scope: their own work, their team, several teams.
Where they fell short of the level
At least one item, even for your strongest person.
Comparison point
One person at the same level, on any team, whose performance
you'd consider similar, and why.
Evidence sources
Your notes, peer feedback, design docs, incident reviews, metrics.
Mark anything you only heard secondhand.
What would change the rating
The fact that, if true, would move it up or down.
The parts that carry the most weight:
- Expectations for the level. Calibration compares people against the level, not against their own last year. An engineer who improved a lot can still be at "meets" for a senior role. Quote your ladder; don't paraphrase it from memory.
- "What changed because of it." "Led the payments migration" is an activity. "Led the payments migration; we retired the old service two months early and on-call pages for that area dropped from weekly to roughly monthly" is an outcome. The room will ask for the second version, so write it now.
- Where they fell short. A case with no weaknesses reads as advocacy and makes everyone in the room discount it. Naming a gap yourself makes the rest of your case more credible.
- The comparison point. This is the question you'll be asked anyway. Picking the comparison yourself, ideally someone from another team, shows you've thought about the shared standard.
If you find you can't fill in the contributions section with specifics, that's the real finding. It usually means you haven't been keeping notes during the period, and the rating is resting on your memory of the last six weeks. Fix that for next cycle with a running document per person, updated after each one-on-one.
Check your own ratings first
Before the meeting, look at your ratings as a set. The biases that calibration is designed to catch are much easier to spot in your own spreadsheet than to defend against in the room.
| Check | What to look for | What to do |
|---|---|---|
| Recency | Most of the evidence is from the last two months | Go back through notes, docs and tickets from the start of the period |
| Visibility | Higher ratings for people on launches and demos; lower for people doing operations, reviews, mentoring | Ask what would have broken without the less visible work |
| Similarity | The people you rate highest work and communicate the way you do | Reread their cases against the ladder wording only |
| Halo | One strong trait (speed, technical depth) lifts every dimension | Rate each ladder dimension separately before the overall |
| Central tendency | Everyone is "meets" because it's the safe answer | For each person, write one sentence on why not one step higher and why not one step lower |
| Distribution | Your team is much more generous or strict than the org's guidance | Not automatically wrong, but you need a reason you can say aloud |
Also check the pattern across groups. If people working part-time, remotely or in a different time zone sit consistently lower, look hard at whether you're rating performance or presence. Your HR partner can help you look at the numbers without singling anyone out.
In the room
A typical calibration runs through people level by level, with each manager presenting briefly and others asking questions. Some habits help:
- Lead with the rating and the strongest evidence, in under a minute. "Proposing exceeds for Ana, senior engineer. The level asks for leading cross-team technical work; she designed the event schema that three teams now publish to and ran the review with all of them." Then stop.
- Answer questions with evidence, not effort. "She worked really hard" and "he's been through a lot this year" are about effort or circumstances. They may matter to how you support the person; they aren't the level.
- Say "I don't know" when you don't. If someone asks about an area you have no evidence on, say so and offer to find out. Guessing in the room is how ratings lose credibility.
- Ask other managers the same questions you'd want asked of you. "What changed because of it?" and "Who's the comparison?" are fair questions for everyone, and asking them consistently is the actual job of the meeting.
- Don't trade. "I'll let your person go to exceeds if you back mine" feels collegial and turns calibration into a negotiation between managers. Each rating stands on its own case.
- Change your mind when the evidence warrants it. If someone shows you a comparison you hadn't considered and it holds up, move the rating and say why. Managers who never move lose influence over time; the ones who move for good reasons are trusted when they hold firm.
When you manage the managers
If you're the one running calibration for your managers, your job is the standard, not the outcome. Before the meeting, share the ladder wording and two or three anonymized example cases at each rating, so everyone is calibrating against the same picture. That is the "concrete scale" idea from Noise applied directly. In the meeting, spend the most time on ratings near a boundary and on any manager whose distribution is far from the rest. Afterwards, tell each manager privately what pattern you saw in their ratings. That feedback is often more useful than the individual changes.
After the meeting
Write down what changed and why while it's fresh: for each moved rating, the evidence or comparison that moved it. You'll need it for the review conversation and for next cycle.
When a rating went down, don't tell the person "calibration lowered you." It shifts the decision to a faceless committee, and it suggests you didn't believe in the rating you're now delivering. Own the result: explain what the level expects, what the evidence showed, and what would lead to a different result next time. If you genuinely disagree with the outcome, raise it with your own manager, not with the person you're reviewing.
Finally, look at the cases that struggled. Most of them will have the same weakness, usually thin evidence from the first half of the period. That tells you what to change in how you take notes and run one-on-ones for the next cycle.
A checklist for the week before
- Each person has a written case of half a page to one page, written before seeing other managers' ratings.
- Every case quotes the ladder for that level and names at least one shortfall.
- Every proposed rating has a comparison point, ideally from another team.
- You've run the bias checks across your whole team, not person by person.
- You know which ratings are near a boundary and what evidence would move each one.
- You can say, in one sentence, why your team's distribution looks the way it does.
If you only do one thing, take your highest and lowest proposed ratings and ask: would I give the same rating if this person had the same results but worked on a different team, for a different manager? If you hesitate, that's the case to rewrite first.
Performance management, calibration and the biases that affect ratings are covered in the strategic team leadership section of the PTMA study guide.