Goals That Stick: OKRs, RAG Status and Honest Reporting
Why team goals die by February and how to fix it: writing real OKRs, a weekly tracking cadence, RAG statuses that stay honest, and ending sandbagging.
Every team has lived this cycle: goals get set in a burst of January optimism, formatted into a slide, presented once — and then nobody looks at them until a panicked week in late March when someone asks how the quarter went. The goals didn't fail because they were wrong. They failed because nothing connected them to the following Tuesday.
Goal systems have two separable parts: the format (OKRs, or whatever your company calls its version) and the operating system — the cadence of tracking, the honesty of status reporting, and the incentives that determine whether people write real goals or safe ones. Almost all goal failure lives in the second part, which is why switching formats never fixes it. This guide covers both: OKRs done plainly, the failure modes that kill most rollouts, RAG status reporting that stays honest, the sandbagging problem and its actual causes, and a weekly-to-quarterly cadence that keeps goals alive for thirteen weeks instead of two.
OKRs in one page
The format, stripped of consultancy varnish:
An Objective is a qualitative statement of what you want to be true by the end of the period. Good objectives are directional and motivating: "New customers succeed without hand-holding." "Our release process is boring." An objective answers where are we going.
Key Results are the two to four measurements that would prove you got there. Each is a number with a current value and a target: "Support tickets per new customer in their first month: from 4.2 to under 2." "Deploys requiring manual intervention: from 40% to under 10%." Key results answer how would we know.
The discipline of the format is the separation: objectives carry meaning, key results carry measurement, and neither tries to do the other's job. A quarter's worth for one team is two or three objectives, each with two to four key results — roughly six to ten numbers total. Teams carrying fifteen key results aren't ambitious; they're unprioritized, and by week four they're tracking none of them.
Two more distinctions that earn their complexity:
- Committed vs. aspirational. A committed KR is a promise — you expect to hit 100%, and missing triggers real replanning. An aspirational KR is a stretch — landing at 70% is a good quarter. Label every KR as one or the other when you write it, because a team that doesn't know which game it's playing will either sandbag everything or miss everything, and both corrode trust in the system.
- Key results are not tasks. "Launch the referral program" is a task — you can complete it while achieving nothing. "Referred signups: from 0 to 150/month" is a key result. Tasks belong on the board as bets you're making to move the KR; if you ship the referral program and the number doesn't move, the correct response is a new bet, not a checkmark. This one confusion — activity dressed as outcome — is the most common defect in real-world OKRs, and it's worth auditing for explicitly: read each KR and ask, "could we 'complete' this and still have failed?" If yes, rewrite it.
Writing key results: a quick quality test
A well-formed KR passes four checks:
| Check | Failing example | Passing rewrite |
|---|---|---|
| Outcome, not activity | "Ship onboarding redesign" | "Week-one activation: 34% → 50%" |
| Measurable with a source you name | "Improve customer happiness" | "NPS from in-app survey: 31 → 40" |
| Influenceable by this team | "Company revenue: +20%" (for a support team) | "Churn from support-flagged accounts: 8% → 5%" |
| Baseline included | "Reach 95% test coverage" (from where?) | "Coverage: 62% → 80%" |
The baseline check looks pedantic and isn't: a target without a stated starting point is unfalsifiable ambition, and hunting down the baseline is frequently where teams discover their metric isn't actually instrumented — better to learn that in planning week than in week eleven.
Pair lagging targets with leading indicators
One refinement pays for itself across the quarter: for any KR that moves slowly, name the faster-moving number you'll watch in the meantime. Churn, NPS, and revenue are lagging indicators — by the time they move, the quarter is mostly spent, which makes weekly reviews of them an exercise in staring at a flat line. A KR like "logo churn: 3.1% → 2.4%" should carry a companion leading indicator the team can influence and read weekly: percentage of at-risk accounts contacted within 48 hours, weekly active usage among renewal-window customers, time-to-resolution on churn-flagged tickets. The lagging number remains the KR — it's the outcome you actually want — but the leading number is what the biweekly review argues about, because it's the only place mid-quarter course correction is possible. Teams that skip this end up with reviews that alternate between "no change" and "too late."
Why OKR rollouts fail
Most organizations that "tried OKRs and they didn't work" hit one of five walls, all avoidable and all predictable.
The cascade waterfall. Leadership decrees that every team's OKRs must mechanically derive from the level above, and planning becomes a six-week alignment bureaucracy where teams reverse-engineer their real work into inherited phrasing. Alignment matters, but it's a conversation, not a spreadsheet join: teams should see company objectives, draft their own contribution, and reconcile in one working session. If planning takes longer than two weeks, the process has become the product.
Set-and-forget. The January slide, the March panic. No tracking cadence means the goals were never operational — this is the failure the entire second half of this guide exists to fix.
Everything is priority one. Five objectives, twenty KRs, and a team of six. The format can't fix a prioritization failure; it can only expose one. If it matters enough to track weekly, it fits in three objectives. The rest is the roadmap, not the quarter.
Goals as performance reviews. The moment OKR scores feed compensation, every rational employee sandbags. This failure mode is important enough to get its own section below.
The format-only adoption. The company adopts the vocabulary — objectives! key results! — while changing nothing about cadence, honesty, or incentives. Six months later OKRs are "just more paperwork," which is accurate, because paperwork is all that was implemented.
RAG status: three colors that have to mean something
Between the quarterly bookends, the workhorse of goal tracking is the humble red/amber/green status on each key result and objective. RAG works — it's glanceable, it aggregates, executives parse it instantly — but only if the colors have operational definitions and honest incentives. Left undefined, every status drifts green until the week it snaps to red, a pattern so common it has a name: watermelon reporting. Green rind, red inside.
Define the colors by trajectory and action, not vibes:
| Status | Operational definition | What the owner must attach |
|---|---|---|
| Green | On pace to hit the target at current trajectory | Nothing — greens get four seconds in review |
| Amber | A named risk makes the target uncertain; recoverable with action now | The risk, and the mitigation being tried |
| Red | Will miss without intervention beyond the team's own control | The decision or help needed, from whom, by when |
Three rules keep the colors truthful:
Amber is a request, not a confession. The entire value of RAG is early warning, and early warnings only come from people who've seen them welcomed. When the first amber of the quarter appears, the manager's response is being watched by everyone: "good catch, what do you need?" produces a team that flags in week three; interrogation produces a team that's green until week twelve. If your statuses are all green all quarter, you don't have a healthy team — you have a reporting system nobody trusts. (This dynamic is the same trust mechanic that underpins accountability without micromanagement.)
Colors come with numbers. A RAG status on a KR that has a current value is checkable: "green, 41% against a 50% target, week eight" invites a useful argument in a way "green" does not. Statuses without the underlying number regress to mood reporting.
Red triggers a decision, not a ceremony. A red that sits red for three weeks means the escalation path is broken. Every red should generate, within a week, one of: added resources, cut scope, a changed target with a written reason, or a conscious decision to accept the miss. Reds that just age are the system telling you reviews have become theater. The craft of writing status updates people actually read — narrative plus RAG, standing questions, history — is covered in depth in Project Status Reports People Actually Read.
Sandbagging: the incentive problem underneath
Sandbagging — setting targets you've secretly already hit, or padding estimates until "stretch" means "safe" — is the most misdiagnosed goal pathology. Managers treat it as a character flaw; it's actually a rational response to two specific incentive errors, and it disappears when the errors are fixed.
Error one: scoring goals like exams. If hitting 100% of your KRs is "success" and 70% is "failure," people will write KRs they can't miss. The fix is to say — and mean, and demonstrate — that for aspirational KRs, landing around 60–70% is the intended outcome, and a team that goes 10-for-10 will be asked to aim higher next quarter, not congratulated. Google's internal guidance famously treated ~0.7 as the healthy landing zone for stretch goals, and the norm travels well. The test of whether you mean it: what happened, publicly, to the last team that set an audacious target and hit 55% of it? If the answer is "nothing good," your sandbagging is manufactured upstream.
Error two: wiring goals to compensation. The moment OKR attainment feeds bonuses or ratings, you've converted a planning tool into a negotiation, and everyone negotiates rationally: lowball the target, hide the upside, celebrate the inevitable "overachievement." Keep the systems separate — performance reviews assess the person (judgment, execution, growth, collaboration) with goal context, not goal arithmetic. Someone who took a genuinely hard swing and landed 60% may have outperformed someone who coasted to 110% of a padded number, and your review process has to be able to say so, or your goal process will lie to you forever.
There's a softer variant worth naming: estimate inflation under interrogation. If every amber gets cross-examined, owners quietly pad every future target to keep statuses green. The padding tax is invisible and compounding. The antidote is the same amber-protection norm as above — which is why incentive design, not format, is the heart of goal systems.
One more incentive note: sandbagging's mirror twin, heroic overcommitment, is just as damaging — the team that commits to double what it can do, delivers half, and calls it culture. Committed KRs exist precisely to prevent this: a committed target is one you'd bet the team's credibility on, and a team should feel the difference every time it labels one.
The cadence: how goals survive thirteen weeks
Everything above is configuration. Cadence is the runtime. Goals stay alive through three nested loops, each cheap, each with a named owner:
| Loop | Cadence | Time | What happens |
|---|---|---|---|
| Update | Weekly | 5 min per KR owner | Owner updates the number, sets the color, writes one line on trajectory |
| Review | Biweekly (weekly late in quarter) | 15–20 min in the team meeting | Skim greens, discuss ambers, decide on reds; adjust bets on the board |
| Score & reset | Quarterly | 90 min | Score each KR 0.0–1.0, run the goal retro, draft next quarter |
Mechanics that make the loops stick:
Every KR has exactly one owner — the person who updates the number and answers for the trajectory. Not the person who does all the work; the person who makes sure the number is current and the risks are voiced. Shared KRs go stale precisely because "the team" owns them.
Statuses live where the team already looks — the same space as the boards and check-ins, not a slide deck in someone's drive. The update should be a two-minute edit, not a document production event. Teams on Openbook typically run this as a Project Status room — goals with RAG status, narrative updates with a history timeline, standing questions for the weekly one-liner — with a Dashboard pulling live numbers from their boards, so the Monday review is "open the room" rather than "who has the latest version."
The review discusses bets, not just numbers. The interesting biweekly question isn't "what's the number" (pre-read) but "is our current work actually moving it?" If activation has been flat for four weeks while the team ships onboarding tweaks, the review is where someone gets to say the tweaks aren't working and propose a different bet. This is the connective tissue between goals and the board — without it, OKRs and day-to-day work drift apart within a month, and you're back to the January slide.
Scoring is boring, and that's the point. At quarter's end, each KR scores as fraction of target achieved (from 0 to 1, capped); the interesting part is the twenty minutes of goal retro after: Which KRs were badly written — activity dressed as outcome, wrong baseline, wrong metric? What did we learn about our capacity? What carries over — and carryover should be a deliberate re-commitment, not a default, because a KR that rolls forward three quarters running is either mis-sized or not actually a priority. Feed the lessons straight into next quarter's draft; the scoring session and the planning session should be the same week. (For running the planning side without burning two days, see Quarterly Planning Without the Two-Day Offsite.)
Transparency: goals are public by default
Publish every team's objectives, key results, and current RAG status where the whole organization can read them. The case is practical, not ideological:
- Dependencies surface early. The platform team discovers in week two — not week eleven — that three teams' KRs assume the new API ships mid-quarter.
- Duplicate work becomes visible. Two teams independently targeting "reduce onboarding friction" find each other in planning week.
- Status honesty compounds. A public amber is mildly uncomfortable once; a public green that everyone can see is false is untenable. Sunlight does quiet enforcement work that no review meeting can.
- People connect their week to something. The engineer deciding between two tasks on a Tuesday can check which one touches a KR. That tiny, frequent act of self-alignment is most of what "alignment" actually means in practice.
The legitimate exceptions are narrow: sensitive personnel matters, unannounced M&A or partnerships, and security work whose details shouldn't circulate. Handle these by abstracting the public version ("Objective: close our top audit findings" with the specifics access-controlled), not by making whole teams' goals private. If more than a sliver of your goal tree is secret, the secrecy is usually habit, not necessity.
What transparency does not mean: leaderboards. Publishing scores ranked by team, or celebrating the "best OKR attainment," reintroduces the exam incentive through the side door and re-manufactures sandbagging at scale. Publish for coordination, never for competition.
A worked example
To make it concrete, here's a realistic quarter for an eight-person product-engineering team, written the way this guide recommends:
Objective 1: New users reach value without our help. (aspirational)
- KR1: Week-one activation rate: 34% → 50% (owner: Dana)
- KR2: Onboarding-tagged support tickets: 120/mo → 60/mo (owner: Chris)
- KR3: Time from signup to first core action: 26 min → 10 min (owner: Dana)
Objective 2: Releases are boring. (committed)
- KR1: Deploys with manual steps: 40% → under 10% (owner: Priya)
- KR2: Change-failure rate: 12% → under 5% (owner: Priya)
Week 6 review, condensed: O1-KR1 is amber — activation is at 38%, trajectory says ~44% by quarter end; Dana's mitigation is killing the tour nobody finishes in favor of a checklist, decision needed on design time. O1-KR2 is green at 84/mo. O1-KR3 is red — the number hasn't moved and the team's two bets both failed; the review decision is to stop polishing signup and test pre-filled workspaces instead. O2 is green across the board. Total meeting time on goals: eighteen minutes. That's the whole system working: numbers current, one honest red, a changed bet, and no ceremony.
At quarter end: activation lands at 46% (score 0.75 — a good aspirational outcome), tickets at 55/mo (1.0), time-to-value at 14 min (0.75 after the pivot), both release KRs hit (1.0, as committed KRs should). The retro notes that KR3 was nearly a task in disguise and that the team's first two bets were guesses; next quarter's KRs each get a named leading indicator.
Getting started: one team, one quarter
You don't need an org-wide mandate. The system proves itself at the single-team level in one quarter:
- This week: Draft two or three objectives with two to four KRs each. Run every KR through the four-check table. Label committed vs. aspirational. Assign one owner per KR and write down every baseline.
- Before the quarter starts: Stand up the status page where the team already works. Put the biweekly 20-minute review on the calendar for the full quarter. Announce the amber-protection rule out loud.
- During: Owners update weekly. Reviews discuss ambers and bets. Reds get decisions within a week. You, the manager, thank the first amber publicly and never let a red age.
- Quarter end: Score, retro the goals themselves, and draft next quarter in the same week — carrying forward lessons, not defaults.
Run that loop once and you'll have something rare: goals that were as true in week eleven as they were in week one, and a team that trusts the colors on the page.
If you want the infrastructure without assembling it, Openbook's Project Status room gives you goals with RAG tracking and a full update history, Dashboards wire the numbers live from your boards, and everything sits in the same space as the work itself. See the features and set up your first quarter's goals this week.