Prove
All outcomes
Teams

Complete a Full Performance Review Cycle

12 weeks · 3 milestones

Run end-to-end reviews: calibration, written assessments, delivery, and follow-through.

Milestone map

Milestone map

3 milestones

Design the Performance Review Framework

2–4 weeks for framework design (before the review period begins or in the first 2 weeks of the review data collection window)

A performance review cycle is an organisational management system: it clarifies expectations (what does good performance look like for this role?), captures evidence of performance over a period (what did this person accomplish, and how?), provides calibration (are assessments consistent across managers?), and produces actionable conclusions (what changes, what stays the same, what does this person need to grow?). The design phase must produce a framework that covers all four components before the cycle begins. Designing the cycle after starting it produces inconsistent assessments and contested decisions.

Proof required

Submit: (1) a performance review framework: the review period, the performance dimensions being assessed (2–4 specific dimensions with operational definitions — e.g. 'Technical Execution: quality and reliability of code shipped to production, measured by incident rate and review round-trips'), the rating scale with definitions for each level, and the calibration process (who reviews ratings for consistency before they are finalised); (2) a template for written performance assessments: the sections each manager fills in for each direct report, including supporting evidence requirements; (3) the communication plan: how and when assessments will be shared with employees, and what follow-up actions are expected after delivery.

What gets checked

  • Performance dimensions have operational definitions — 'communication' is not an operational dimension; 'Communication: clarity of written and verbal updates in project meetings and async channels, as assessed by cross-functional partners' is operational; definitions must be specific enough that two managers would assess the same person consistently
  • Rating scale has level descriptions, not just labels — 'Exceeds Expectations' is a label; 'Exceeds Expectations: delivers the stated outputs with consistently lower defect rates than the team average AND actively identifies and resolves adjacent problems not originally in scope' is a level description
  • Calibration process names participants and decision rules — 'managers will calibrate' is not a process; 'all people managers will attend a 90-minute calibration session where each manager presents their highest and lowest ratings with evidence, and any rating that differs from peer assessment by more than one level requires written justification' is a calibration process

Common mistakes

  • Performance dimensions that assess personality rather than observable outputs — 'attitude', 'culture fit', or 'enthusiasm' are not observable outputs; every dimension must be assessable from observable work evidence
  • Calibration without a decision rule — a calibration meeting where managers share their ratings and leave unchanged is a peer-pressure session, not a calibration; the calibration process must specify how discrepancies are resolved
  • Communication plan that lacks a follow-up action requirement — delivering performance assessments without an agreed follow-up leaves the assessment as feedback rather than a management action; the plan must specify what each employee does with their assessment

Resources

Foundationstart here

Depthgo deeper

Masteryfor the dedicated

What a verifier looks for

  • Dimension operationality: ask the submitter to demonstrate how they would assess one dimension for a hypothetical employee who did good work but communicated poorly — does the dimension definition make the assessment unambiguous?
  • Calibration decision rule: ask what happens when a manager gives a 'Meets Expectations' rating and the calibration group believes the person deserves 'Below Expectations' — is there a defined decision rule?
  • Follow-up requirement: ask the submitter what each employee is expected to do after receiving their assessment — if the answer is 'they can choose what to do with it', the communication plan lacks a follow-up action requirement

You'll sign in first, then come straight back here.

Run the Full Review Cycle: Calibration, Delivery, Follow-Through

4–8 weeks to run the full cycle (assessment writing, calibration, delivery, follow-up commitment)

The review cycle runs end-to-end: managers complete written assessments for each direct report, the calibration session is held and discrepancies are resolved, assessments are delivered to employees with follow-up action plans, and growth commitments are recorded. A performance review cycle that stops at delivery — where assessments are shared but no follow-up actions are committed to or tracked — is incomplete. The follow-through is the management system; without it, the review is a bureaucratic exercise.

Proof required

Submit: (1) anonymised written assessment samples for at least 3 employees (different performance levels): each sample must show the performance dimension ratings with evidence, the supporting examples, and the follow-up actions; (2) a calibration log: who attended, how many ratings were changed after calibration and in which direction, and one specific example of a rating that was adjusted with the reason; (3) a follow-through tracker: for all employees in the review cycle, the agreed follow-up action (growth area, new project, role change, or no change) and the timeline for the first check-in.

What gets checked

  • Written assessment samples show evidence, not conclusions — '4/5 on Technical Execution' without supporting examples is a conclusion; '4/5 on Technical Execution: shipped 3 features without production incidents over the past 6 months; the auth module redesign reduced incident rate from 3/month to 0' is evidence-based
  • Calibration log shows real changes — if no ratings changed in calibration, either the framework is working perfectly (unlikely on a first cycle) or calibration was nominal; the log must show at least 1 rating change with the reason
  • Follow-through tracker covers all employees — a tracker that covers only employees who received difficult feedback is not complete; the tracker must cover all employees with their follow-up action, including 'no change planned' for employees already performing well

Common mistakes

  • Written assessments with no evidence — an assessment that says 'strong communicator' without a supporting example is a judgement, not an evidence-based assessment; every rating must have at least 1 supporting example
  • Calibration that changes ratings downward without evidence — if a manager's rating is changed down in calibration without a specific evidence-based reason, the calibration is being used to manage ratings to a curve rather than to ensure consistency; the calibration log must show the reason for each change
  • Follow-through actions that are vague aspirations — 'focus on communication skills' is not an action; 'attend 2 cross-functional planning meetings in the next quarter as a presenter, with feedback from the PM lead by week 8' is an action

Resources

Foundationstart here

Depthgo deeper

What a verifier looks for

  • Evidence quality: ask the submitter to read the supporting evidence for the highest and lowest ratings in one written assessment — are they specific enough to be verified against work records?
  • Calibration authenticity: ask the submitter to describe the most challenging calibration discussion — what was the disagreement, and how was it resolved?
  • Follow-through specificity: ask the submitter to read the follow-through action for the employee with the clearest growth need — is it specific enough to be assessed at the next check-in?

You'll sign in first, then come straight back here.

Review First-Cycle Outcomes and Improve the Framework with Expert Review

1–2 weeks for retrospective, framework improvement plan, and expert review (at 8 weeks post-delivery of assessments)

The final milestone requires a retrospective on the first complete review cycle, a lessons-learned document, and a real-time review with a named, qualified reviewer. The ADVERSARIAL VERIFICATION Level 2 standard applies: the reviewer must challenge the framework design (were the dimensions predictive of performance?), the calibration quality (did calibration increase or decrease consistency?), and the follow-through rate (what percentage of employees completed their first check-in action?).

Proof required

Submit: (1) a cycle retrospective: framework coverage (which dimensions were most and least useful), calibration consistency (did the first calibration session reveal systematic biases in any manager's assessments?), delivery reception (how did employees respond to their assessments — any surprises, contested ratings?), and follow-through rate at 8 weeks (what percentage of follow-through actions were completed or in progress?); (2) a framework improvement plan (300 words minimum): what you would change in the dimension definitions, rating scale, calibration process, or follow-through mechanism for the next cycle; (3) documentation of a real-time Q&A session with a named reviewer who has ≥3 years of people management experience and a verifiable professional profile — the documentation must include specific challenge questions and your responses; the reviewer must not be a direct report.

What gets checked

  • Cycle retrospective covers the follow-through rate at 8 weeks — a review cycle whose follow-through rate is unknown 8 weeks later has not been tracked; the retrospective must report the fraction of follow-through actions that were completed or in progress
  • Framework improvement plan addresses the hardest dimension to calibrate consistently — almost no first-run framework has all dimensions working well; the improvement plan must name the dimension that produced the most calibration difficulty and propose a specific revision
  • Q&A shows the reviewer challenged the follow-through rate — a reviewer who asked only about the assessment delivery process without asking 'what happened to the growth commitments after the review?' did not meet the adversarial standard

Common mistakes

  • Retrospective that only covers the assessment delivery — performance reviews are only valuable if the follow-through actions produce change; a retrospective that covers the review process but not the follow-through outcomes is incomplete
  • Framework improvement plan with no specific revision — 'we will improve the rating scale' is not a revision; 'we will replace the Communication dimension with Cross-Functional Impact (defined as: quality of communication in situations where alignment with other teams affects project delivery, assessed by cross-functional partner feedback)' is a revision
  • Reviewer who participated in the review cycle — a reviewer who was assessed in the cycle, managed someone in the cycle, or helped design the framework cannot give an independent adversarial view

Resources

Foundationstart here

Depthgo deeper

What a verifier looks for

  • Follow-through rate: ask the submitter to report the fraction of employees who completed or were actively working on their follow-up action at 8 weeks — if the answer is unknown, the cycle was not followed through
  • Hardest dimension: ask the submitter which dimension produced the most disagreement in calibration — a confident answer with an example shows genuine retrospective analysis
  • Framework improvement specificity: ask the submitter to read the proposed revision for the hardest-to-calibrate dimension — is it specific enough to produce more consistent assessments in the next cycle?

You'll sign in first, then come straight back here.

We use analytics to improve Powstik. No ads, ever.