Blog

Preparing for calibration so your ratings hold up

How to build a one-page case for each person, spot your own rating biases before others do, and argue fairly in a calibration meeting without trading favors.

Calibration is the meeting where your ratings meet everyone else's. You walk in with "exceeds" for the engineer who carried the migration, someone across the table asks "compared to what?", and you find yourself telling a story about how hard they worked. Ten minutes later the rating is "meets", and you have to explain that to a person who trusted you to represent them. Most ratings that get moved in calibration don't get moved because they were wrong. They get moved because the manager couldn't show why they were right. This post covers how to prepare a case that survives the room, how to check your own ratings first, and how to behave once you're in there.

What calibration is for

Calibration exists because managers rate differently. Some are generous, some are strict, and most rate people who resemble them a little higher than people who don't. The meeting is meant to apply one standard across teams, so that "exceeds" means the same thing whichever manager you happen to report to.

That problem is bigger than most managers assume. Steven Scullen, Michael Mount and Maynard Goff studied multi-rater performance ratings of several thousand managers (published in the Journal of Applied Psychology in 2000) and found that the rater's own idiosyncratic tendencies explained more of the variation in ratings than the performance of the person being rated. Daniel Kahneman, Olivier Sibony and Cass Sunstein devote a chapter of Noise (2021) to performance ratings and come to a similar conclusion: much of what a rating measures is the rater. One of their remedies is to make the scale concrete, anchoring ratings to shared descriptions and examples rather than to adjectives like "strong" that each manager reads differently.

So treat calibration as a check on your judgment, not an attack on your team. If you walk in assuming your ratings are accurate and everyone else's are inflated, you'll argue badly and learn nothing.

What calibration should not be is forced ranking with a quota of low ratings. If your company runs a hard distribution, you still calibrate against the level expectations first; the distribution conversation comes after, and you should know which of the two you're having.

Build the case before the meeting

Each person you're rating needs a short written case. One page at most; half a page is better. Write it before you see anyone else's ratings, so you're anchored on evidence rather than on what other managers proposed.

A template:

Name, level, proposed rating

Expectations for this level (one line, quoted from your career ladder)

Top three contributions this period
  For each: what they did, what changed because of it, and how you know.
  Name the scope: their own work, their team, several teams.

Where they fell short of the level
  At least one item, even for your strongest person.

Comparison point
  One person at the same level, on any team, whose performance
  you'd consider similar, and why.

Evidence sources
  Your notes, peer feedback, design docs, incident reviews, metrics.
  Mark anything you only heard secondhand.

What would change the rating
  The fact that, if true, would move it up or down.

The parts that carry the most weight:

  • Expectations for the level. Calibration compares people against the level, not against their own last year. An engineer who improved a lot can still be at "meets" for a senior role. Quote your ladder; don't paraphrase it from memory.
  • "What changed because of it." "Led the payments migration" is an activity. "Led the payments migration; we retired the old service two months early and on-call pages for that area dropped from weekly to roughly monthly" is an outcome. The room will ask for the second version, so write it now.
  • Where they fell short. A case with no weaknesses reads as advocacy and makes everyone in the room discount it. Naming a gap yourself makes the rest of your case more credible.
  • The comparison point. This is the question you'll be asked anyway. Picking the comparison yourself, ideally someone from another team, shows you've thought about the shared standard.

If you find you can't fill in the contributions section with specifics, that's the real finding. It usually means you haven't been keeping notes during the period, and the rating is resting on your memory of the last six weeks. Fix that for next cycle with a running document per person, updated after each one-on-one.

Check your own ratings first

Before the meeting, look at your ratings as a set. The biases that calibration is designed to catch are much easier to spot in your own spreadsheet than to defend against in the room.

Check What to look for What to do
Recency Most of the evidence is from the last two months Go back through notes, docs and tickets from the start of the period
Visibility Higher ratings for people on launches and demos; lower for people doing operations, reviews, mentoring Ask what would have broken without the less visible work
Similarity The people you rate highest work and communicate the way you do Reread their cases against the ladder wording only
Halo One strong trait (speed, technical depth) lifts every dimension Rate each ladder dimension separately before the overall
Central tendency Everyone is "meets" because it's the safe answer For each person, write one sentence on why not one step higher and why not one step lower
Distribution Your team is much more generous or strict than the org's guidance Not automatically wrong, but you need a reason you can say aloud

Also check the pattern across groups. If people working part-time, remotely or in a different time zone sit consistently lower, look hard at whether you're rating performance or presence. Your HR partner can help you look at the numbers without singling anyone out.

In the room

A typical calibration runs through people level by level, with each manager presenting briefly and others asking questions. Some habits help:

  • Lead with the rating and the strongest evidence, in under a minute. "Proposing exceeds for Ana, senior engineer. The level asks for leading cross-team technical work; she designed the event schema that three teams now publish to and ran the review with all of them." Then stop.
  • Answer questions with evidence, not effort. "She worked really hard" and "he's been through a lot this year" are about effort or circumstances. They may matter to how you support the person; they aren't the level.
  • Say "I don't know" when you don't. If someone asks about an area you have no evidence on, say so and offer to find out. Guessing in the room is how ratings lose credibility.
  • Ask other managers the same questions you'd want asked of you. "What changed because of it?" and "Who's the comparison?" are fair questions for everyone, and asking them consistently is the actual job of the meeting.
  • Don't trade. "I'll let your person go to exceeds if you back mine" feels collegial and turns calibration into a negotiation between managers. Each rating stands on its own case.
  • Change your mind when the evidence warrants it. If someone shows you a comparison you hadn't considered and it holds up, move the rating and say why. Managers who never move lose influence over time; the ones who move for good reasons are trusted when they hold firm.

When you manage the managers

If you're the one running calibration for your managers, your job is the standard, not the outcome. Before the meeting, share the ladder wording and two or three anonymized example cases at each rating, so everyone is calibrating against the same picture. That is the "concrete scale" idea from Noise applied directly. In the meeting, spend the most time on ratings near a boundary and on any manager whose distribution is far from the rest. Afterwards, tell each manager privately what pattern you saw in their ratings. That feedback is often more useful than the individual changes.

After the meeting

Write down what changed and why while it's fresh: for each moved rating, the evidence or comparison that moved it. You'll need it for the review conversation and for next cycle.

When a rating went down, don't tell the person "calibration lowered you." It shifts the decision to a faceless committee, and it suggests you didn't believe in the rating you're now delivering. Own the result: explain what the level expects, what the evidence showed, and what would lead to a different result next time. If you genuinely disagree with the outcome, raise it with your own manager, not with the person you're reviewing.

Finally, look at the cases that struggled. Most of them will have the same weakness, usually thin evidence from the first half of the period. That tells you what to change in how you take notes and run one-on-ones for the next cycle.

A checklist for the week before

  1. Each person has a written case of half a page to one page, written before seeing other managers' ratings.
  2. Every case quotes the ladder for that level and names at least one shortfall.
  3. Every proposed rating has a comparison point, ideally from another team.
  4. You've run the bias checks across your whole team, not person by person.
  5. You know which ratings are near a boundary and what evidence would move each one.
  6. You can say, in one sentence, why your team's distribution looks the way it does.

If you only do one thing, take your highest and lowest proposed ratings and ask: would I give the same rating if this person had the same results but worked on a different team, for a different manager? If you hesitate, that's the case to rewrite first.

Performance management, calibration and the biases that affect ratings are covered in the strategic team leadership section of the PTMA study guide.

Working toward PTMA Associate?

The free study guide covers every competency area on the PTMA exam, with practical examples.

More from the blog

Practical articles for technical managers at every stage. All posts

  1. Delegating work without dropping it or hovering over it

    Decide how much authority to hand over, brief the work in five lines, and set check-ins that match the person's experience, not your nerves.

  2. Research bets that end in a decision instead of fading out

    How to frame exploratory research as a bet with the hardest question first and written kill criteria, so every project ends in a clear go or stop.

  3. A one-page technical strategy your executives will read

    Most engineering strategies are lists of goals. How to write one that makes real choices, fits on a page and gets a decision from the people who fund it.