The evidence

Why this works

We built this on a specific behavioural mechanism, and we think you should be able to check our work — including the parts that argue against us.

Everything below is cited. Where the research is contested, we say so. Where it doesn't support us, we say that too.

01

The mechanism everyone gets backwards

Behavioural psychology splits into four quadrants that people routinely collapse into two. Positive and negative describe whether something is added or removed. Reinforcement and punishment describe whether the behaviour goes up or down.

Add somethingRemove something
Behaviour increases Positive reinforcementBadges, streak confetti, praise. Negative reinforcementThe aversive thing stops when you act. The seatbelt chime. This is us.
Behaviour decreases Positive punishmentInsults after a failure. What people assume an app called yourbully does. Negative punishmentLosing something you already had.

Nearly every "tough love" app reaches for positive punishment — make you feel bad after you fail. It's the weakest quadrant for this purpose. Punishment delivered after the fact doesn't teach the behaviour you want; it teaches escape from whatever delivers it. In an app, the available escape is the uninstall button.

Negative reinforcement is different, and it's unusually durable — avoidance learning is famously resistant to extinction, because the behaviour keeps being rewarded by the absence of the thing you're avoiding.

So every aversive element here is pending, announced, and cancellable by the behaviour. A countdown you can stop. A stake that burns at midnight unless you act. You are never punished for who you are — you're handed something you can switch off, and switching it off is the behaviour we want.

02

Whether stakes actually work

Voluntary commitment contracts — where you put something of your own at risk — have held up in randomised trials across several domains.

Smoking

Giné, Karlan and Zinman tested a product called CARES with smokers in the Philippines. Participants deposited their own money for six months, then took a urine test for nicotine and cotinine. Pass, money back. Fail, money to charity.

+3pp Smokers offered CARES were three percentage points more likely to pass the test at six months than controls — and the effect persisted in surprise tests at twelve months, after the contract had ended.

Weight

Volpp and colleagues randomised 57 participants across monthly weigh-ins, a lottery incentive, and a deposit contract.

14 lb Mean additional weight lost by the deposit-contract group over 16 weeks compared with control. 47.4% hit the 16-pound target, against 10.5% of controls.

How the money is framed

Patel and colleagues found that incentives allocated upfront and then removed on failure outperformed gain-framed incentives and lotteries for physical activity. That's why we show the full stake in your account on day one and subtract from it visibly, rather than holding a deposit you hope to earn back.

Worth knowing this isn't universal: a later trial in university students found loss framing performed worse than gain framing and reduced goal commitment. We treat it as a default worth testing, not a settled result.

03

Why we make you write an if-then plan

The gap between intending something and doing it is well documented, and one of the most reliable fixes is embarrassingly simple: specify in advance when, where and how. "If [situation], then I will [action]."

d = 0.65 Gollwitzer and Sheeran's meta-analysis of 94 independent tests found a medium-to-large effect on goal attainment. The effect on preventing derailment of an effort already underway was larger still, at d = 0.77.

This is why the app refuses "exercise more." It makes you name the trigger and the response, and it pushes you to put the action ahead of the moment you usually fail — "change into gym clothes before I sit down," not "go to the gym." That single field does more work than anything else in onboarding.

04

Why we will never call you fat

This is the question we get most, so here is the direct answer: because it would make the product worse at its job.

Shame and guilt are not interchangeable. Tangney and Dearing's framework separates shame — focused on the self, on what you are — from guilt, focused on the behaviour, on what you did. Because the self feels harder to change than an action, shame predicts withdrawal, concealment and hostility. Guilt predicts repair.

Shame — about the self"You're lazy and you'll always be fat."
Guilt — about the behaviour"Third Thursday in a row. Same excuse window."

And in the specific domain people expect us to be cruellest about, the outcome data is unambiguous. Sutin and Terracciano followed 6,157 participants in the Health and Retirement Study across four years.

Worse Participants who reported weight discrimination were more likely to become obese, and more likely to remain obese, than those who didn't. The same group later found weight discrimination associated with roughly 60% higher mortality risk across two large cohorts, not explained by common physical and psychological risk factors.

Insults don't work. They make the thing worse. So we don't sell them — and the alternative is harder to argue with anyway, because it's specific, true, and about something you can change tonight.

05

Why we post misses, not goals

The common version of social accountability is announcing your goal publicly. That version is contradicted by the research.

Gollwitzer, Sheeran, Michalski and Seifert ran four experiments and found that identity-related intentions noticed by other people were acted on less intensively than intentions that went unnoticed. Social recognition appears to deliver a premature sense of already being the person you're trying to become. The effect held among participants strongly committed to the goal.

We post misses, never intentions. A missed commitment doesn't grant an identity symbol — it withdraws one. That runs in the opposite direction from the effect that undermines goal-announcement products.

You write the exact wording in advance, while calm. There's a fifteen-minute cancel window before anything sends, a hard cap of one post a week, and it's disabled entirely for sensitive categories.

06

Why squads are capped at twelve

When everyone's outcome depends on the group's performance — an interdependent group contingency — behaviour changes measurably. The Good Behavior Game is the most studied version.

TauU .82 Effect size across 137 phase contrasts covering 1,580 students. In a head-to-head comparison, most children preferred the interdependent arrangement to an individual one.

But it only works while your contribution is visible. In a group of ten thousand, one person moves the number by a rounding error, and a contingency you can't influence stops being a contingency at all. That's why the research uses teams, and why we cap squads at twelve and switch the mechanic off below four.

Squad rewards are gift codes and cosmetics — never money. Nobody's stake is linked to anybody else's week.

07

What the evidence doesn't support

You'd read this from someone else eventually. Better from us.

This app has not been clinically trialled

The studies above tested commitment contracts and if-then planning. They did not test our product. We are applying findings, not reporting our own.

Effects fade when the contract ends

Volpp's group found substantial weight regain after the incentive period stopped, and follow-up work by John and colleagues found the same pattern. Commitment devices reliably produce behaviour during the commitment. Maintaining it afterwards is genuinely unsolved and we don't pretend otherwise.

Uptake is low even when it works

Only 11% of smokers offered CARES took it. Products like ours work for people who want this kind of pressure. That's a real limit on who we can help, not a marketing problem to engineer around.

The shame/guilt distinction has critics

Some researchers argue the boundary is blurrier than the standard account allows, and have raised methodological objections to the instruments used to measure the two. We think the design implication survives the critique — target behaviour, not identity — but the underlying science is contested and we won't overstate it.

Loss framing isn't universally better

It outperformed gain framing for physical activity in Patel's trial and underperformed it in a later student trial. We default to loss framing and treat it as a hypothesis we're testing, not a fact.

We are not treatment

If you're dealing with an eating disorder, a substance dependence, or a mental health condition, this app is not the intervention you need. Certain goal types are locked in our product for exactly that reason.

08

The obvious questions

So you profit when people fail?

We never insult users — that rule is published and enforced by an automated check before any message sends. But on the money: yes, we charge you when you miss, and that is our revenue. We'd rather state it plainly than have you discover it.

The conflict is real, so here is how it's constrained. The core product is free — we don't charge a subscription on top. Your first week is $0. Then $5, and the amount only rises after you have already missed: $5, $10, $20, $40, $75, $150, $300, $600, $1,000. You set a cap on every goal, nothing takes more than $1,000 from you in a month, and one word ends everything permanently at no cost. The doubling schedule exists so you reach a motivating amount fast without paying anything at amounts too small to motivate — and so nobody pays real money before the product has already given them something.

What if someone with an eating disorder signs up?

We screen at signup. Food, body and exercise goals accept actions only — never weights, calories or targets — and a positive screen caps those categories at timers and reminders. Exercise is included deliberately: driven training is part of the same pattern often enough that stakes are the wrong tool. We'd rather lose that customer than hurt them.

Why call it yourbully, then?

Because everyone knows what it means, and nobody knows what "accountability platform" means. The name describes how it feels at 11pm when you haven't done the thing. It doesn't describe what we say to you — every rule about that is published, and enforced by an automated check before any message sends.

Why not send the money to charity instead?

We tried to design that and the law made it a bad idea: routing user money to a third party is money transmission, which needs licensing in nearly every state. Keeping the fee ourselves is ordinary merchant activity, which is why every long-running product in this category works this way.

There's also a behavioural argument. Money going somewhere good softens the miss — it turns a loss into a donation, which is exactly the feeling that stops the stake from working. Paying us has no upside at all, and that's the point.

Can I quit?

Instantly, free, from any screen. You set a kill phrase when you write your contract; typing it ends everything. It ends the contract — it doesn't pause it. There's no pause, no vacation mode, no snooze and no grace period, because those are why other apps are decorative. Quitting is always available. Drifting isn't.

Why won't you build a summer-body mode?

Because it's an outcome goal in the one domain where outcome goals do measurable harm, dressed as a seasonal promotion. We build seasonal skins — they change the voice and the look, never the rules — and that's one we've decided not to make.

09

References

  1. Giné, X., Karlan, D. & Zinman, J. (2010). Put Your Money Where Your Butt Is: A Commitment Contract for Smoking Cessation. American Economic Journal: Applied Economics, 2(4), 213–235. doi:10.1257/app.2.4.213
  2. Volpp, K. G. et al. (2008). Financial Incentive-Based Approaches for Weight Loss: A Randomized Trial. JAMA, 300(22), 2631–2637. doi:10.1001/jama.2008.804
  3. John, L. K. et al. (2011). Financial Incentives for Extended Weight Loss: A Randomized, Controlled Trial. Journal of General Internal Medicine, 26, 621–626.
  4. Patel, M. S. et al. (2016). Framing Financial Incentives to Increase Physical Activity Among Overweight and Obese Adults. Annals of Internal Medicine, 164(6), 385–394.
  5. Gollwitzer, P. M. & Sheeran, P. (2006). Implementation Intentions and Goal Achievement: A Meta-Analysis of Effects and Processes. Advances in Experimental Social Psychology, 38, 69–119.
  6. Gollwitzer, P. M., Sheeran, P., Michalski, V. & Seifert, A. E. (2009). When Intentions Go Public: Does Social Reality Widen the Intention-Behavior Gap? Psychological Science, 20(5), 612–618.
  7. Tangney, J. P., Stuewig, J. & Mashek, D. J. (2007). Moral Emotions and Moral Behavior. Annual Review of Psychology, 58, 345–372.
  8. Tangney, J. P. & Dearing, R. L. (2002). Shame and Guilt. Guilford Press.
  9. Sutin, A. R. & Terracciano, A. (2013). Perceived Weight Discrimination and Obesity. PLoS ONE, 8(7), e70048.
  10. Sutin, A. R., Stephan, Y. & Terracciano, A. (2015). Weight Discrimination and Risk of Mortality. Psychological Science, 26(11), 1803–1811.
  11. Ashraf, N., Karlan, D. & Yin, W. (2006). Tying Odysseus to the Mast: Evidence from a Commitment Savings Product in the Philippines. Quarterly Journal of Economics, 121(2), 635–672.
  12. Good Behavior Game meta-analysis: TauU = .82 across 137 phase contrasts, 1,580 students.
Try the loop

Ninety seconds, one real contract

Write one real contract, watch a window close, and decide for yourself whether the pressure is the useful kind.

Start