Strategy

North Star Metrics and the Numbers That Quietly Mislead

One headline metric, a handful of inputs, and a working knowledge of how each can be gamed, misread, or ruined by its denominator.

A North Star metric is one number that captures the value customers get from your product, chosen so that moving it reliably grows the business. The point is not measurement. It is alignment: a shared answer to "is this worth doing?" that does not require re-litigating strategy every time.

What makes a good one

Four tests. A candidate that fails any of them will cause problems within two quarters.

  1. It reflects customer value. Something good happened for a user, not just for you.
  2. It leads revenue. It moves before the money does, so it works as a steering signal.
  3. The team can influence it. Share price cannot be a North Star. Neither, usually, can raw signups.
  4. It is hard to game without doing real work. This is the one people skip, and it is the one that bites.
ProductWeak choiceBetter choice
Collaboration toolRegistered usersWeekly active teams with 3+ contributors
MarketplaceListings createdSuccessful transactions per week
Learning productVideos startedLessons completed to assessment
Analytics productDashboards createdDashboards viewed 3+ times in a week

Every weak choice is easy to move without helping anyone. You can drive signups with a giveaway and listings with a bulk import. The better versions all require the customer to actually get something.

Input metrics matter more day to day

The North Star is too slow and too far away to guide weekly work. Break it into three or four inputs the team can move directly.

If the North Star is weekly active teams with 3+ contributors:

  • Invitations sent per new workspace
  • Invitation acceptance rate
  • Time from signup to first shared document
  • Week-two return rate for invited members

Now a sprint can target something specific. "Improve invitation acceptance from 34% to 45%" is work you can plan. "Increase weekly active teams" is not.

Five ways numbers mislead

1. Goodhart's law

When a measure becomes a target, it stops being a good measure. Target support ticket closure time and tickets get closed prematurely. Target feature adoption and you get modals nagging people into clicking once. The metric improves; the product gets worse.

Guard: pair every target with a counter-metric. Closure time paired with reopen rate. Adoption paired with retention of adopters.

2. Averages hiding the distribution

Average session length rose from 4 to 6 minutes. Good news, unless it happened because casual users left entirely and only power users remain — the average rose while the business shrank.

Guard: look at medians and percentiles. Averages conceal bimodal populations.

3. The denominator moving

Activation rate jumped 8 points the week marketing paused a low-quality acquisition channel. Nothing about the product changed; the mix of people entering did.

Guard: when a rate moves, check the numerator and denominator separately before believing it.

4. Survivorship bias

Surveying current users tells you what people who stayed think. The ones who left — the group holding your most valuable information — are silent by construction.

Guard: churn interviews. Painful, consistently the highest-signal research available.

5. Cohorts versus snapshots

Total active users can rise for a year while every individual cohort retains worse than the last, because acquisition is outrunning a leak. Snapshots hide it; cohorts expose it immediately.

Guard: retention by signup cohort, always. It is the single most honest chart in product analytics.

Vanity metrics have a tell

Ask what you would do differently if the number halved. If there is no answer, it is not a metric, it is a mood. Cumulative totals are the classic case: they only ever go up, so they can never tell you anything is wrong.

Frameworks worth knowing

  • AARRR (pirate metrics) — acquisition, activation, retention, referral, revenue. Best for finding where the funnel leaks.
  • HEART — happiness, engagement, adoption, retention, task success. Google's model, strongest for UX-quality questions.
  • Leading vs lagging — not a framework so much as the habit that matters most. Revenue is lagging. Activation is leading. Steer with leading, report with lagging.

A workable setup

  1. One North Star, reviewed quarterly, changed rarely.
  2. Three or four input metrics the team owns and reviews weekly.
  3. A counter-metric for anything actively targeted.
  4. A cohort retention chart nobody is allowed to remove.
  5. A written note of what each metric excludes — internal accounts, trials, bots. Six months on, nobody remembers, and someone will make a decision on a number that never meant what they think.

Metrics tell you whether the outcome moved. Choosing which outcome to chase is prioritisation, and communicating it is the roadmap.

Keep reading