I've been spending a lot of time with Peter Grafe lately. He runs BlueAlpha, a company I advise, and he has the same problem with dashboards that I picked up at Affirm: he doesn't believe them.
One of his clients is MUBI, the global film streaming service, distributor, and producer. They release titles in cinemas and on the platform across dozens of markets on a near-weekly cadence, which makes them a hard business to measure. When MUBI ran their own numbers, they found that in a single US week, Meta alone had driven roughly thirteen times the signups its dashboard took credit for. Paid and promotions together were more than half their growth in normal weeks, not the 20 to 30% they had planned paid around.

I have looked at a lot of budgets and the same gap is in most of them. The difference is that MUBI went looking.
At Affirm, we were working toward GAAP profitability. If you have never worked inside that, it is hard to describe what it does to a marketing team (also, I’m sorry for the PTSD). Every discretionary dollar gets questioned. Marketing is the largest, softest, most discretionary line on the page, so it is the first place finance looks when the model has to close. We could not protect the budget by telling a good story. We either proved the spend produced incremental margin, or we gave it back so the number could get to black.
Nobody upstairs cared what Meta or Google claimed we drove. What mattered was whether a dollar of spend produced a dollar of incremental contribution, or whether we had paid to stand in front of people who were coming anyway. When the whole company is trying to close the gap to profitability, "coming anyway" is the difference between a channel that earns its keep and a channel you are wasting money on while telling the board you are being disciplined.
So we ran holdouts. Geo tests, ghost ads, before-and-afters where we shut a channel off in half the country and watched what happened to signups and volume. We did it because we had no other way to keep the budget. Once you watch a "winning" channel fall apart under a real test, you read dashboards differently.
Branded search was the one that got me. It looked incredible in the platform. Huge return, a clean last-click story. Then you hold it out and most of that volume shows up anyway, because those were people typing our name who had already decided. We were charging ourselves a toll to harvest demand we had already created, and booking the toll as growth. Meanwhile the upper-funnel work that was creating the demand, and never earns a last click, got starved because attribution refused to credit it.
Years later at Webflow, one of the first things I did was shut paid branded search off completely. Self-serve acquisition moved less than 3%. We had been paying to catch people who were already coming. So we pulled that capital out and pushed it upmarket, toward demand we had to go create. The dashboard had been calling branded a top performer the whole time…it’s wild to me how many orgs are still doing this. I know I’ll hear the conquesting and brand reputation counter, but I seldom buy it.
Attribution and incrementality are two different numbers, and when there is real money on the line, the gap between them is where you either find your margin or lose it without noticing.
BlueAlpha exists to take what I learned the slow way and turn it into a system a marketing team can run on every week, without a month of manual analysis. MUBI is the cleanest proof of it I have seen. Before walking through it, it helps to see how the problem gets built in the first place.
The trap almost every marketing team is standing in
You plan the budget bottom-up. A target per product, per market, per campaign, each line sized by an assumption of how much of that outcome your spend is responsible for. It all rolls up into the plan you get judged on. And every one of those assumptions rests on attribution, which is wrong in a consistent direction.
Attribution error does not wash out at scale the way random noise would. Platforms overcredit the last click and undercredit everything that built the demand upstream, every time. So when you allocate on attribution, you make the same mistake every quarter: you overfund harvesting and underfund creation. You pour money into the channels that close demand, starve the channels that make it, then read the dashboard, see the closers winning, and pour in more.
Run that loop for three years and you have not just wasted budget. You have under-grown the business, bent the whole company toward short-term harvesting, and produced a spreadsheet that certifies you were careful the entire time. Nobody involved was careless. They were careful about the wrong number.
You end up running the budget like a project manager, justifying the last decision and gathering data to guess at the next one, when the job was always to run it like a portfolio manager, moving money to where it demonstrably pays off. Nearly everyone defaults to project mode because you cannot run a portfolio on numbers you do not trust, and attribution never gives you numbers you can trust.
How it showed up at MUBI
MUBI's plan was built exactly the way I just described. Targets are set per title, they roll up into market-level goals, and the budget to hit them is sized by an assumption of how much of each title's signups paid media is responsible for.
The assumption was that paid drove 20 to 30% of new trials. It rested on the same broken input everyone fights, but MUBI has a particularly nasty version of it. A curation-led business that cannot follow members cleanly after signup ends up proxying acquisition with something like the first title watched. Someone clicks an ad for one film, signs up, and watches a completely different one. The proxy breaks, and the platforms stack claimed conversions on top of it.
Every business has its own version of this. At Affirm the click was a checkout and the outcome was a loan that either performed or did not. The proxy always breaks somewhere. The question is whether you have built a way to see past it, or whether you are budgeting on it anyway.
Rory Japp, MUBI's VP of Digital Marketing, described their side of it this way:
"We knew paid was doing more than ad platforms were giving credit for. We just couldn't prove it. BlueAlpha built and validated the models across our global business in four weeks, faster than we could have hired for it, let alone built it."
That first half is where almost every growth team I talk to is standing. Knowing and not being able to prove it is the reason you cannot get the money.
So MUBI built a way to see past it. BlueAlpha shipped a new marketing mix model every week, trained on MUBI's own back-end data. For one US week, the model put Meta's real, incremental number at roughly thirteen times what the dashboard reported. And across normal, non-release stretches, paid's true contribution was well above the 20 to 30% in the plan. Paid media and promotions together had crossed half, at 51% and rising.
Because the plan understated what paid drove, MUBI underfunded it. On Meta specifically, the team estimated they were spending about 20% of what they should have been, on a channel that had once driven one of their strongest acquisition periods.
The cost was never just the budget line
A wrong number at the top does not stay in the budget. It sets how the team spends its days.
When you cannot trust the read, you react. MUBI worked week to week, making promo calls inside tight windows, because they could not see far beyond the week they were in. When the report finance scrutinizes has to be built by hand, your sharpest people spend their hours producing numbers instead of acting on them. And when a channel cannot be measured cleanly, you freeze it. Connected TV, the channel they most wanted to scale, sat untouched for exactly that reason.
That costs more than any single mispriced channel. A whole team stuck reacting and defending when their job is to allocate and grow. You cannot place a bold bet on a read you do not believe, and you cannot move fast when every call has to be relitigated by hand.
What BlueAlpha actually built, judged the way I judge any vendor
I have sat through enough measurement pitches to have a reflex. I look for the three places these products die: the model is something nobody on the team trusts, the read never leaves the report, and it gets ripped out the second the champion changes jobs. Most tools in this category fail all three.
On trust: MUBI's model runs on their own back-end actuals, not platform data, and every read from the model gets checked against an incrementality test. The model produces the read and the test confirms it. Then it retrains weekly, because a model that is right in January and stale by March will steer you wrong just as reliably as no model at all. That is what lets you move real money on it.
On the read leaving the report: this is where nearly every measurement tool I have seen fails. It produces a number, the number goes in a deck, and nobody acts on it. BlueAlpha wired the measurement into the signals that tell you when to move (creative fatigue, pacing, competitive shifts) and then into execution, so a decision becomes a change in the Meta or Google or TikTok account, not a recommendation somebody has to remember to action on a Thursday. That gap between knowing and doing is the part the industry keeps underbuilding, and it decides whether measurement changes anything.
On surviving contact with reality: it deploys into MUBI's own accounts and workspace and gets fed the context that makes their calls theirs: the brand, the history, the constraints, the reasoning behind past decisions. So the output shows up inside their world instead of as generic best practice, and it improves the longer it runs. That is why this kind of build sticks where point tools and agencies get dropped.
Two more things I will credit. The rollout is sequenced so it compounds instead of overwhelming: a trustworthy read and self-generating reporting first, to hand the team back its time; then recommendations they review together; then execution, widened only once they believe it. And the humans keep the wheel. The system does the work, the people own every call. Most builders in this space chase autonomy because it demos well. What MUBI got is a team that can move quickly on numbers it believes.
What the first real read revealed
For the first time, MUBI could see what its budget actually does. Averaged across the full period, the North America model put every line of spend on the same axis: roughly 53% organic, 31% paid, 16% promotions. The top paid channel came in at nearly a 5x return, delivering 54% of paid's contribution on 46% of its spend, which is what a channel looks like when it is both the most efficient thing you own and the most starved. And it came with its own health metrics: 93.8% average three-week-ahead backtesting accuracy, clean convergence, 12 quality checks passing, a full weekly retrain.
Do not skip that last part. Backtesting means the model was scored on data it had never seen, so the read was predictive, not just plausible. A number you cannot inspect is a number finance is right to dismiss.
Three things changed in the same motion. Paid stopped being a defensible guess and became a case for funding the winners instead of spreading thin across all of them. The monthly report started generating itself, handing back the days the team used to lose assembling it. And CTV, the bet that had been too fuzzy to fund, moved onto a structured incrementality-testing roadmap. Going from "we cannot measure it so we will not touch it" to "here is the test that will tell us" is most of what this job is.
North America is only the start. The same system extends across every market MUBI operates in, and up the whole budget, past digital media into brand, promotions, and theatrical, the parts of the mix that have always been the hardest to measure and therefore the easiest to over- or under-fund badly.
The lesson runs past film, and past paid
Strip out the film business and the pattern is everywhere. If your budget is planned bottom-up on assumed contributions that rest on platform attribution, you cannot defend the number, and you are almost certainly underfunding what works while calling it discipline.
The fix is measuring what genuinely drives growth and running the whole budget like a portfolio, where every line, including the lines that never had a dashboard, gets a price on the same axis. You cannot run a portfolio when a third of your positions have no quoted price. That is what a marketing mix model built on your own data buys you: a real price on spend you had no read on before. Brand, promotions, and theatrical sit next to paid on one axis, and you can choose between them instead of guessing.
Paid is where the proof is cheapest to stand up first, because it has the cleanest feedback loop. Once that read holds, the same logic runs up the chain into the spend that has always hidden from measurement. The discipline that reaching profitability forced on me at Affirm is the discipline every growth org will eventually be held to, because the moment money gets expensive, "trust us" stops being an acceptable answer and "here is what each dollar drives" becomes the only one.
MUBI could not answer that question until they built the system that let them. Then paid went from the line they had to apologize for to the case they could make with evidence.
So here is the question Peter would put to you, and the one I had to learn to answer before anyone would let me spend real money:
If finance asked you tomorrow to prove what each dollar of your budget actually drives, not what your platforms claim, could you?