All outcomes
Teams

Scale an Engineering Team from 5 to 20

24 weeks · 5 milestones

Scale an engineering team from 5 to 20 people — the hardest management transition in software. The skills that made you effective at 5 will actively harm you at 20. Every milestone here is about replacing personal relationships with systems, and replacing your presence with documented authority.

Milestone map

Milestone map

5 milestones

Diagnose what breaks when you add headcount

2–3 weeks. The interviews take one week. Writing the document takes another. The most important bottleneck is usually the one that takes longest to name — build in time to sit with the findings before writing.

Before hiring anyone new, audit your current team's systems for what will break at 2x headcount. Interview every current team member individually — ask what slows them down, what information they cannot find, and what decisions they have to wait for. Write a scaling risk document that names the top five bottlenecks ranked by severity. This document becomes your hiring and process roadmap.

Proof required

Share your scaling risk document (minimum 500 words). It must name at least five specific bottlenecks with: what the current state is, what breaks when the team doubles, and what the proposed fix is. Share anonymised notes from your team interviews — at least one direct quote per person. Write 200 words on the bottleneck that surprised you most and why you did not see it before the audit.

What gets checked

  • Five bottlenecks are named specifically — not 'communication could be better' but 'decisions about API design require my sign-off and I am the bottleneck for 6 engineers; at 12 engineers this blocks two squads simultaneously'
  • Interview notes contain at least one direct quote per team member — paraphrases do not count; the actual words reveal what the team member was willing to say, not what the manager heard
  • Proposed fix for each bottleneck is a system, not an event — not 'improve communication' but 'create an architecture decision record (ADR) process so API decisions are documented and do not require my real-time input'

Common mistakes

  • Skipping the team interviews and writing the bottleneck document from your own observations — managers consistently miss their own bottlenecks because they are inside them; the interviews exist to find what you cannot see
  • Ranking bottlenecks by how easy they are to fix rather than by severity — the hardest bottleneck to fix is usually the most important one; a document that lists easy wins first is optimised for comfort, not for scale

Resources

Foundationstart here

Depthgo deeper

Masteryfor the dedicated

What a verifier looks for

  • Read the five bottlenecks. Are they specific enough to hand to a new hire as context for what they are joining? Vague bottlenecks ('process needs work') are not actionable. Named bottlenecks ('the on-call rotation breaks when a third team is added because our alerting is not team-scoped') are actionable.
  • Check the interview notes. Is there at least one direct quote per person? Ask the submitter to read one quote that surprised them. A manager who cannot point to a surprising quote did not listen — they conducted a survey.
  • Ask: 'Which bottleneck on your list is your fault?' A manager who can name a bottleneck they personally created (usually a decision gate or a process gap) is doing the right kind of audit. A manager who lists only structural or process bottlenecks has not looked hard enough at themselves.
  • Check the proposed fixes. Are they systems (processes, tools, structures) or events (meetings, conversations, decisions)? At scale, systems outlast events. A fix that requires the manager's ongoing involvement is not a fix.

Design the team structure for 20 engineers

2–3 weeks. The design takes a week. The team presentation and incorporating feedback takes another. The design will change based on the presentation — budget time for a second draft.

Design the organisational structure your team will have at 20 engineers — before you hire anyone new. This means: how many sub-teams, what each sub-team owns, how sub-teams communicate, how decisions are made between sub-teams, and the manager-to-engineer ratio at each level. Write this as a team design document, then present it to your current team and document their reactions — the concerns they raise now are the problems you will face at scale.

Proof required

Share your team design document (diagrams and text, minimum 600 words). It must include: the org chart at 20 engineers with named roles (not names), the ownership boundaries of each sub-team, the decision-making process between sub-teams, and the communication cadence. Share anonymised notes from the team presentation including at least three specific concerns the team raised. Write 200 words on how you resolved or plan to resolve the most serious concern.

What gets checked

  • Org chart shows roles, not names — a document that only works for the current team is not a scaling design; the structure must work for people who have not joined yet
  • Ownership boundaries are specific — each sub-team owns named systems, services, or domains; overlap must be explicitly named as a known tension with a documented resolution process
  • Three team concerns are specific objections, not vague worries — 'what if sub-teams do not communicate' is vague; 'the current shared infrastructure work will fall into a gap between sub-team A and sub-team B with no clear owner' is specific

Common mistakes

  • Designing the org chart around current people rather than current needs — a team structure built to fit existing personalities rather than to own existing systems will break when any of those people leave
  • Not presenting it to the team before finalising — the people doing the work know where the ownership gaps are; skipping the presentation step means inheriting their known concerns as your future incidents

Resources

Foundationstart here

Depthgo deeper

Masteryfor the dedicated

What a verifier looks for

  • Read the org chart. Can you identify what each sub-team owns without asking the submitter? If the ownership is ambiguous from the document alone, the design has gaps that will become incidents.
  • Check the overlap areas. Does the document name them explicitly? Gaps in ownership are more dangerous than clear ownership of difficult problems. Ask: 'Who owns the deployment pipeline in your design?' If the answer is 'shared' without a resolution process, it will be owned by nobody.
  • Read the three team concerns. Ask the submitter to describe what happened in the room when the most serious concern was raised. Their description reveals whether they ran a genuine consultation or a show presentation.
  • Ask: 'What will you do if a new VP joins in 6 months and wants to reorganise the team differently?' A manager who has thought about structural resilience will have an answer. A manager who built the structure around themselves will struggle.

Hire five engineers without lowering the bar

12–20 weeks depending on the hiring market and role. Five hires at an average 4-week process each. Running processes in parallel is required — sequential hiring takes too long at this stage.

Run five engineering hiring processes end-to-end — from job description to offer signed. Document every candidate who made it to final round and the decision with a specific reason. For every offer extended, document what you compromised on (if anything) and why. For every offer declined or rejected, document the deciding factor. This hiring log is your proof that you maintained the bar rather than rationalised exceptions.

Proof required

Share your hiring log (anonymised — Candidate A through Z). It must show: all candidates who reached final round, the decision, and the specific reason in one sentence each. For the five who received offers, show the deciding factor. Write 250 words on how your definition of a strong candidate changed or did not change across the five hires — did the bar drift? If so, when and why?

What gets checked

  • Hiring log covers all final-round candidates, not just the ones who received offers — a log that only shows successful hires is not a log, it is a list of employees
  • Each decision has a specific reason — not 'not the right fit' but 'strong on systems design but could not articulate trade-offs in the debugging exercise; at our current stage we need engineers who can debug independently'
  • The bar reflection is honest — if the bar drifted on hire 4 or 5 due to time pressure, it must be named as drift, not reframed as learning what we actually need

Common mistakes

  • Hiring for culture fit without defining what culture fit means in measurable terms — 'we liked them' is a rationalisation; the hiring log forces a specific reason, which prevents this failure mode if you write the reason before the offer goes out
  • Letting urgency lower the bar — the fifth hire is the most dangerous one; the team is larger, the pressure to fill is higher, and the rationalisation is easier; the log's bar reflection must address whether this happened

Resources

Foundationstart here

Depthgo deeper

Masteryfor the dedicated

What a verifier looks for

  • Count the candidates in the hiring log who did NOT receive offers. If there are zero, the log is incomplete — at least some final-round candidates should not have received offers. A hiring process with a 100% offer rate has either a very narrow funnel or has not been selective.
  • Read the decision reasons. Are they specific enough to replicate? A reason like 'strong on distributed systems, weak on communication in the design session — not the right stage for us' is replicable. 'Not quite right' is not.
  • Ask the submitter: 'If you ran this process again tomorrow for the same role, what would you change?' A manager who improved the process across five hires has a specific answer. A manager who ran the same process five times without adjusting has not been learning.
  • Check the bar reflection. If the bar drifted, is it named honestly? Ask: 'Was hire 5 as strong as hire 1?' The honest answer to this question reveals more about the hiring process than any other question.

Build a team that works when you are absent

2 weeks of absence plus 1 week of preparation and 1 week of documentation on return. The preparation week is not optional — a team that is not briefed on the authority matrix before you leave will default to waiting for you regardless of what the document says.

Take two consecutive weeks completely off — no Slack, no email, no code reviews, no decisions. Before you leave, document every decision type your team faces regularly and who is authorised to make each one without you. After you return, document every decision that was made in your absence, every decision that was delayed waiting for you, and every incident that occurred. The ratio of made-without-you to delayed-for-you is your delegation score.

Proof required

Share the decision authority matrix you wrote before leaving (a table: decision type, who is authorised, what the escalation path is if that person is also unavailable). Share your return document: decisions made in your absence (with who made them), decisions delayed for you (with why they waited), and any incidents with their resolution timeline. Write 200 words on your delegation score — what percentage of decisions were made without you, and what the delayed ones reveal about where your team still depends on you.

What gets checked

  • Two consecutive weeks, not two separate weeks — the second week is always harder than the first; most delegation failures surface in week 2 when the team has exhausted the decisions they felt comfortable making
  • Decision authority matrix covers at least ten distinct decision types — not just 'technical decisions' but named categories like 'change the on-call rotation', 'approve a vendor contract under $X', 'merge a breaking change'
  • Return document names every delayed decision specifically with the reason it waited — 'we were not sure if we could approve this' is a gap in the authority matrix, not a reason

Common mistakes

  • Checking in during the absence — even one Slack message resets the team's expectation that you are available; the milestone requires complete absence, which means telling the team explicitly that you will not respond and meaning it
  • Writing the authority matrix so broadly that everything is delegated without real authority — 'the team can make any decision' is not an authority matrix; each decision type needs a named person or role who owns it

Resources

Foundationstart here

Depthgo deeper

Masteryfor the dedicated

What a verifier looks for

  • Verify the absence was two consecutive weeks. Ask for evidence — an out-of-office reply, a Slack status showing unavailable, or a team member who can confirm. The milestone specifically requires consecutive weeks.
  • Count the decision types in the authority matrix. If there are fewer than ten distinct types, the matrix is too coarse. Ask: 'Who decides whether to merge a breaking change at 5pm on a Friday?' If the answer is 'the team' without a named person, the matrix has a gap.
  • Read the delayed decisions. For each one, ask: 'Why did this decision wait for you?' The answer should point to a specific gap in the authority matrix — either the decision type was not covered, or the person who was authorised did not feel confident using the authority.
  • Ask: 'What would you change in the authority matrix before your next absence?' A manager who can name specific gaps from the return document has done the post-mortem correctly. A manager who says 'it worked fine' has not read the delayed decisions carefully enough.

Deliver a major project through two sub-teams

8–16 weeks depending on project scope. This milestone cannot be manufactured — it requires a real project with real cross-team complexity. If no such project exists in your current roadmap, that is information about whether your team structure is working.

Take a project that requires work from at least two of your sub-teams and deliver it to production without managing the day-to-day execution yourself. Your role is: setting the goal and success criteria at the start, removing blockers that cross sub-team boundaries, and reviewing the outcome at the end. Everything in between is your sub-team leads' work. Document your interventions — every time you got pulled into the execution, write down why and what you would build to prevent needing to be pulled in again.

Proof required

Write a project retrospective (500–700 words) covering: the original goal and success criteria, what each sub-team owned, the three most significant cross-team blockers and how they were resolved, your list of interventions (times you got pulled into execution) with the system you built or changed to prevent each recurrence, and whether the project hit its success criteria. Get one of your sub-team leads to endorse this proof on Powstik — they should confirm your account of how the project ran.

What gets checked

  • Intervention list names each intervention specifically and the system change that followed — not 'I stepped in a few times' but 'I had to resolve a data contract dispute between sub-teams A and B; I subsequently wrote a data contract template that both teams now sign before any inter-team dependency is built'
  • Sub-team lead endorsement on Powstik is required before submission — a manager's account of their own delegation that is not verified by a team member is T0 self-reported; T3 requires the endorsement
  • Success criteria were defined at the start, not after completion — verified by the retrospective's opening section; retrospectives that define success after completion are post-hoc rationalisation

Common mistakes

  • Running a project that only involves one sub-team and calling it cross-team — the project must genuinely require work from two sub-teams with a real dependency between them; a project where one team 'supports' another by answering questions is not a cross-team project
  • Documenting interventions after the project rather than during it — the intervention list must be written in real-time; a list written at the end from memory will miss the interventions that felt normal at the time but are the most important ones to fix

Resources

Foundationstart here

Depthgo deeper

Masteryfor the dedicated

What a verifier looks for

  • Read the intervention list. Is each intervention accompanied by a system change? An intervention list with no system changes is a list of times the manager managed — not a list of times the manager improved the system. The point of documenting interventions is to eliminate them.
  • Check the sub-team lead endorsement on Powstik. Is it from someone who had direct visibility into how the project ran? Ask the endorser: 'Did the manager stay out of the execution?' Their answer should match the intervention list.
  • Read the success criteria defined at the start. Were they measurable? 'Ship the feature' is not a success criterion. 'Ship the feature with p95 latency under 200ms and zero data loss incidents in the first 30 days' is a success criterion.
  • Ask: 'What would you do differently in the next cross-team project?' A manager who lists specific process changes (not just 'communicate better') has learned from the intervention list. A manager who says 'it went well' has not interrogated the interventions carefully enough.

We use analytics to improve Powstik. No ads, ever.