At a glance
- Of 54 four-week gym programs, about 45% significantly raised weekly visits, by 9% to 27%.
- Only 8% still showed a significant, measurable effect after the four weeks ended.
- The top program rewarded members for coming back after a missed workout, and expert forecasters failed to predict the winners.
In this essay
It’s 6 a.m., your gym bag is by the door, and your phone holds a note called “habit tricks.” Lay out your clothes the night before. Tell a friend. Keep the streak alive. Save a favorite podcast for the treadmill. Many of these ideas trace back to a study, and each one sounded convincing on its own.
The common mistake is assuming every trick works, and keeps working. You rarely see how these ideas do when they compete on the same people at the same time. You almost never see what’s left once the push stops.
One team of researchers set up exactly that contest: 54 programs, 61,293 gym members, one scoreboard. By the end, you’ll know which kinds of nudges held up in that head-to-head test, and why you should expect most of them to be small and to fade.
One kitchen, many recipes
Most habit advice is tested the way a family tries a new chili recipe. One pot, one kitchen, one night. Everyone says it’s good. What nobody learns is how it stacks up against the dozens of other recipes online, cooked on the same stove for the same people.
A megastudy is the cook-off. Many ideas run at once, on one population, against one shared comparison group. Every idea meets the same people, in the same season, with the same measuring stick. A trick can’t look strong just because it was tested at a lucky moment or on an unusually eager group.
Sounds like a technicality? It changes the question. The tips you collect usually arrive one at a time, each with its own study. A cook-off asks something harder: not “Does this work?” but “Does this work better than the other things I could try?”
What happened at 24 Hour Fitness
Researchers led by Katherine Milkman and Angela Duckworth ran this kind of cook-off with 24 Hour Fitness, an American gym chain. Their paper appeared in Nature in December 2021. The full paper sits behind a login, so the details here come from summaries published by the authors’ own institutions.
Thirty scientists from 15 US universities worked in small independent teams. Together they designed 54 different four-week digital programs, all aimed at getting members to the gym. All of them ran at the same time on 61,293 members and were measured against the same control group.
Here is the scoreboard. According to Harvard Kennedy School’s summary, about 45% of the programs significantly increased weekly gym visits, with effects ranging from 9% to 27%. Carnegie Mellon’s account counts 53 experimental arms against one control and reports that 24 of them produced a significant increase. Twenty-four divided by 53 is about 45%, so the two line up.
Then the four weeks ended. Only 8% of the programs produced a change that was still significant and measurable afterward. For the rest, any measurable effect was gone once the support stopped.

Which programs came out on top? The single most effective one gave members a small bonus — 125 points, worth about $0.09 — for returning to the gym after missing a workout they had scheduled. That bonus sat on top of a standard reward of a quarter per workout. The second-best program’s twist was information: it told members that most Americans exercise regularly and that exercise rates are trending upward, as Penn Today reported.
The last result is the most humbling. Before anything was known, impartial experts and practitioners forecast which programs would work best.
Forecasts by impartial judges failed to predict which interventions would be most effective.
A later commentary on that forecasting exercise, in the Timmerman Report, found no robust correlation between the predicted effects and the measured ones. One way to read this: on this question, even the professionals were guessing — much as the rest of us do when we pick a habit trick.
Three things that held up
First: reward the comeback, not the streak
The strongest program in the test aimed at the day after a miss, not the first day. Picture it: your Tuesday workout disappears because a meeting ran late. Streak thinking says the chain is broken, and Wednesday starts to feel optional too. The winning design treated that Wednesday as its own small event, with a prize for walking back in.
It suggests how important it is to avoid having a series of missteps when you’re pursuing a goal.
Why would nine cents matter to anyone? As money, it hardly seems worth it. One way to read the result is that the bonus worked more as a signal: coming back counts. What changes is where your attention goes. A missed day becomes a planned decision point instead of a verdict on your character.
Second: treat your own forecast as a guess
If impartial judges couldn’t pick the winners from 54 options, your gut probably can’t pick from the handful on your list. Say you’re sure an alarm across the room will get you up, and that a workout partner won’t help. Those are hunches, not results. What changes is how you hold them. You try one idea for a set stretch, count your actual visits each week, and keep the idea with the higher count — not the one with the better story.
Third: budget for the fade
Of the programs that raised visits, most no longer showed a measurable effect once they ended, so plan for the morning after the support stops. A four-week challenge wraps up. The app stops sending reminders, the points dry up, and next Monday looks like every Monday before it. What changes is timing. Before the last week, you decide what replaces the scaffolding: a new cue, a fresh small reward, or another round of the same plan. The study didn’t test what works after the fade, so this part is your own experiment.
The other side
The study ran for four weeks, and that limits what it can say about habits. Gretchen Chapman, a Carnegie Mellon behavioral economist and one of the co-authors, suggested the window may not have been long enough to solidify a new habit for the long term. So “faded” might partly mean “never had time to set.” The design can’t separate the two.
It also measured one narrow behavior: checking in at a gym. Not eating, not sleep, not saving money. The pattern of small, short-lived effects is a finding about gym-visit nudges. It may carry over to other habits, or it may not. This study doesn’t tell you.
Does that make the whole exercise pointless? Not really. Small is not the same as worthless. According to Penn Today, most of the programs could be deployed at scale for about $0.75 per person per month. A modest bump in visits, spread across a large membership, can add up for very little money. Part of the point of a megastudy is to find which small effects are real and worth using widely, not to dismiss them for being small.
Try this today
Pick one habit you’re already attempting — the gym, a morning walk, anything with a schedule. Then do three small things, on paper.
- Write down the next time you plan to do it, with a day and an hour.
- Write a comeback rule for the day after a miss: “If I skip, I go back the next day, even for a short session, and I get ___.” Keep the prize small. The top program’s was about nine cents.
- Next to the rule, write your guess of how much it will help. Later, check that guess against a count of what you actually did, not against a feeling.
Try it for one day. Not forever. One day. The point isn’t to become a new person by tonight. It’s to have a plan waiting for the morning you miss.
The megastudy didn’t hand anyone a trick that works for everyone. What it offers is better expectations. In a head-to-head test on 61,293 members, the nudges that did best rewarded coming back after a miss, or told people that exercise is normal. Most of the rest moved visits a little or not at all, and most effects faded when the programs ended. Knowing that, a slow week isn’t proof that you failed. In this study, small effects that faded were the norm, not the exception.
Plan for the day after you miss, not just the day you start.




