Somewhere in the life of every shared-house app, somebody adds a leaderboard. It looks obviously correct. Make the work visible, rank the contributions, let the numbers settle an argument that has been running for years.

Then a share of couples uninstall it, and one of them writes a review explaining that their partner refused to have it on their phone because it had turned the housework into a scoreboard.

The folk explanation is roughly backwards

The usual story is that a leaderboard demoralises whoever sits at the bottom. Hydari and colleagues tested that in Management Science in 2023 with 516 participants.

The bottom quartile gained roughly 1,365 steps a day. The top quartile dropped roughly 631. The people it hurt were the ones already winning.

What the reward research measured

Deci, Koestner and Ryan pooled 128 studies in 1999 and ran rewards directly against task interest. For tasks people already found interesting the effect was strongly negative. For tasks people found dull it was slightly positive and statistically insignificant.

Nobody is intrinsically motivated to descale a kettle. Chores sit squarely in the second group, so quoting overjustification at chore points applies a result well outside the conditions it was measured in. This is the part most people get wrong, in both directions.

The cell that does the damage

One arrangement in that analysis stands out: performance-contingent rewards where most participants receive less than the maximum. That produced the largest negative effect in the whole set.

It also describes a ranked leaderboard precisely. Everybody except the person in first place is being told, repeatedly and in public, that they came in under the line.

Points are close to neutral. Position is what does the damage.

What we built instead

  • Points are minutes. A three-tier easy, medium, hard scale cannot represent the gap between flipping a switch and an hour of shovelling, and a real reviewer named exactly that as the reason their spouse would not use a competitor.
  • They only go up. Nothing is deducted, ever, and no streak can be lost.
  • Two clocks. This week resets on Sunday so nobody is permanently last. All time never resets, so the effort stacks.
  • Capacity is adjustable. A household member working sixty hours and a household member on sick leave are not measured against the same expectation.
  • No rank exists. There is no field in the data for position, so no screen in the app can render one.

The rule that holds it together. The numbers can render. Beatrice can never speak them. "The app says you have done nothing this week" is the sentence that kills a household product, and a test asserts she has no way to construct it.

The review that explains the stake

We read 9,064 reviews across 50 apps. The one that stayed with us was five stars, left approvingly. The reviewer had reinstalled a chore app to show their daughter how much they and their husband do compared to her lack of contributing.

They loved the product. They were using it to build a case against somebody in their own family. Software that makes that easy has taken a side in an argument it does not understand, and the household is the thing that pays.