Podcast: Demystifying open source MMM (with Michael Kaminsky)

On this week’s episode of the podcast, I am joined by Michael Kaminsky, co-founder and co-CEO of Recast. We explore the rise of open-source Marketing Mix Modeling (MMM) tools and the challenges of measuring modern ad spend. Among other things, we discuss:
- Whether open-source MMM libraries like Robyn and Meridian are truly unbiased tools or subtle instruments for big tech self-grading
- How marketers can effectively evaluate the accuracy of MMM outputs through techniques like parameter recovery and predictive forecasting
- Why smaller, performance-driven brands might actually find more value in last-touch attribution than in complex econometric modeling
- If the inherent uncertainty in MMM estimates makes them fundamentally incompatible with the fast-paced feedback loops of digital advertising
- What role generative AI and large language models play in democratizing access to sophisticated marketing data science analysis
- When a brand should transition from simple attribution methods to a multi-layered approach involving incrementality and econometric modeling
- How the shift toward CTV and non-trackable channels is forcing a resurgence in probabilistic measurement frameworks like MMM
Thanks to the sponsors of this week’s episode of the Mobile Dev Memo podcast:
- INCRMNTAL. True attribution measures incrementality, always on.
- Xsolla. With the Xsolla Web Shop, you can create a direct storefront, cut fees down to as low as 5%, and keep players engaged with bundles, rewards, and analytics.
- Branch. Branch is an AI-powered MMP, connecting every paid, owned, and organic touchpoint so growth teams can see exactly where to put their dollars to bring users in the door and keep them coming back
Interested in sponsoring the Mobile Dev Memo podcast? Contact Mobile Dev Memo advertising.
The Mobile Dev Memo podcast is available on:
Transcript
Eric Seufert: Welcome to the Mobile Dev Memo podcast. I am your host, Eric Seufert, and I am joined today by Michael Kaminsky. Michael, welcome back to the podcast.
Michael Kaminsky: Eric, thanks for having me. Happy to be here.
ES: The occasion for us speaking, although I had been intending to bring you back on, was an article that I shared in the Mobile Dev Memo Slack recently that was questioning whether open-source MMMs, and in particular Robyn and Meridian, were allowing Meta and Google to grade their own homework. This catalyzed a conversation, and I thought it would be interesting to get your take on that, but also to talk all things open measurement. Before we get to that, could you please reintroduce yourself to the audience?
MK: I am co-founder and co-CEO of Recast. We do marketing measurement, planning, and analysis, including media mix modeling. I have a background in statistics, causal inference, and econometrics, but have been doing marketing science and measurement for the last 10 or 15 years.
ES: Let’s talk about Robyn and Meridian because that was the source of this article. What does it mean for an MMM to be open source? Why do you think some people view Robyn or Meridian as an attempt by Meta or Google to grade their own homework, and why is that view flawed? That was the point I made in the Slack chat. It is a flawed view in my mind because there is no baked-in preference for Meta or Google. There is no way for these open-source libraries to purposefully and deliberately artificially inflate the performance metrics of Meta and Google. Why do you think that view persists, and what does it mean for an MMM to be open source?
MK: I will take the questions one at a time. Why does the view persist? It is reasonable and justified for people to be skeptical when these sorts of ad platforms propose measurement methods. They are obviously deeply and broadly incentivized to try to make their own platforms look good. That is very reasonable to me. I do not think it applies in this case, but if you log into Meta’s ad platform, you should always be asking yourself how their reporting is biased and why it is biased. Meta and Google are obviously very incentivized to make their own ads look good. That should be baked into how you are looking at the reports coming out of any of these platforms.
People who do not understand how open-source software works or how media mix models work probably just apply that same skepticism without thinking deeply about what is actually going on under the hood. Being open source means that all of the underlying code is available for anyone to look at, read, and, in the case of both Meridian and Robyn, download, change, use, and even resell based on the open-source license they have.
I can download all of the underlying source code for both Robyn and Meridian. I can read every single line of it, change that source code, run it on my own data, and see how the results come out. Meta and Google wanted to open source these tools because they wanted to give that extra layer of transparency to the world to be able to say we are not biasing this toward our own channel. You can look at the code under the hood and confirm that for yourself. That is not something that you are able to do with their default ad platform reporting based on lookback windows or even based on their incrementality-driven conversion optimization. We cannot see any of that underlying source code, but for Robyn and Meridian, we can.
ES: That is what strikes me as so odd about making this particular accusation because it is entirely verifiable. By nature of being open source, you can look through everything. If you were going to make that claim, would you not make it against the opaque platform attribution tools? Why would you make that claim against the open-source library where all the code is available to you to vet?
MK: I fully agree. It gets eyeballs and gets people talking. This just feels like a thing that someone wanted to publish to get attention, which works pretty well in the attention economy of today.
ES: I want to hover on two aspects of this. The code is visible to you. There are no starting weights here. It is not like an open-weights model. An open-weights model like a trained model has weights applied to it. This is just the model framework, and so you train it on your own data. There are no starting weights that you might apply to your own analysis out of the box. This is a model that you run on your data. It is meant to be trained on your data to start with. There is no way for them to bias or inflate the performance baseline of their own channels because that is not how the models work. They are a blank slate. It is just the code.
MK: That is exactly right.
ES: There is no starting point. There are no priors here. It is just meant to be run on your data. There is nothing to it until you run it on your data. You could just clone the repo and tell Codex to look through it and see if there is anything in here that would bias Meta performance. That would be a very sensible thing to do. You should do that with any repo to look for vulnerabilities. That would take five minutes. This is not some big mystery. We are not necessarily arguing in the abstract. You could just do that.
MK: I would recommend that people do that. There are other very simple tests that you can run. One that I often suggest is to run the model with Facebook activity labeled and Google activity labeled. Run it once, and then run it again and just switch the names. You will see that it is no different. It is not keyed off of the names or off of anything special about any of the different marketing channels. You can split the marketing channels any way you want, give them all different names, and the results will not change. It is very unclear to me how you would imagine that there could be some inherent bias in these models when you can do these very easy tests to demonstrate that is not true.
ES: I just said the word bias, but I have been trying to not say the word bias because there are two interpretations of that. You could think of bias as being some inherent prejudice against something, which would be more like a cognitive flaw. Or you could think about bias in the context of machine learning, like the bias-variance tradeoff. Bias meaning the risk of underfitting and variance meaning the risk of overfitting. There, I think you could make the case. Here is where I will contradict myself a little bit, but it is not about Meta or Google per se. It is a reason I do think why they did release these as open source.
I think that absent this kind of framework, absent a probabilistic attribution solution, you might actually underfund a Meta or a Google. They might get less credit than they otherwise would, and you might be over-crediting things like linear TV, radio, or out-of-home. If you are primarily a digital property and you are doing those legacy channels, you might actually be under-attributing these primary channels that are driving a lot of top-of-the-funnel awareness, even if they are not necessarily capturing every click. I think that was their motivation here. It was that they think you are under-counting and under-valuing their contribution. If we give you this tool, it will give you the ability to see that actually Meta or Google is generating more value for you than what your measurement apparatus is telling you at the moment. Is that a reasonable assertion, or is that a weird thing to say?
MK: I think it is a little bit more complex than that. Let me talk through what I think is actually going on here. Meta and Google are great advertising products, probably the best ever. What Meta and Google found was that a lot of their advertisers were using other types of media mix models, either ones that they were managing in-house or ones that were run by other big consulting firms or even their agencies.
A lot of those other media mix models are in fact, or have been historically, structurally biased against digital channels. They would do things to put their thumb on the scale for the model to make specifically channels like TV look good. There are a variety of reasons for that. A lot of agencies historically made a lot of their money specifically from TV advertising, so that was a cash cow for them and they wanted to make it look good. A lot of CMOs and even higher-ups in marketing at the brands themselves felt strongly that they wanted to do a lot of brand advertising because it is fun and more interesting than a lot of the conversion-based advertising.
Facebook and Google felt like these MMM results keep coming back saying that Facebook and Google are really bad. We feel quite strongly that they are good, and we think it is because the MMMs are biased against us. They wanted to release these open-source MMM products as a counterweight to what they were seeing coming out of these other MMMs that in fact were structurally biased against digital channels because the analyst or whoever was managing them was putting their thumb on the scale for the channels that they wanted to look good for a bunch of different reasons not actually related to performance. When I think about why they are doing this, it is more about that reason—combating other MMMs which are actually of even lower quality—versus necessarily trying to give a better way of measuring versus a digital tracking type measurement.
ES: Right, I think that I would be much more inclined to trust one of these tools than to trust the output of an agency’s tool because the agency really is incentivized to have you think a certain way because they are taking a cut of spend as their fee.
MK: I think that is exactly right. There are lots of problems with these open-source MMMs and I am happy to talk about what those problems are, but versus some random regression that some analyst at a random consulting firm or a random agency is running, the open-source MMMs are going to be way better than that. They are more inspectable and more interpretable, especially in the world of AI and LLMs. Have Codex or Claude look at it, have it pressure test it, and that is just going to be a much better experience for most brands.
ES: Walk me through what you think the flaws and shortcomings are of these open-source MMMs.
MK: The default assumption should be that MMMs in general are wrong. The problem that an MMM is trying to solve—looking at aggregate data and variations in that data and then trying to attribute causality to that—is just an almost impossibly difficult problem. It is so difficult and there are so many ways for it to go wrong. Easy ways to demonstrate this to yourself are to simulate data that you think is realistic based on how you believe marketing works in the real world. Run that through any MMM and then see if you get the correct results back that matched your simulation.
When you make very obvious assumptions in the way that you simulate the data—for example, that marketing performance can change over time as creative changes or as the market changes, that there is a lot of correlation between your different marketing channels, that there is seasonality—you run the data that you simulated so you know what the truth is. You know truly how effective every marketing channel in the mix is. You run it through the MMMs and you get results back that do not match that at all.
Obviously on its face, the MMM is not working. It is not able to find the true causal signal in that data. That is the fundamental problem with MMMs. Over the last couple of years, people got really excited about marketing mix models, especially because these open-source tools were released. As people are using them now today, the shine is starting to wear off because people run their data through the model one week and then they run it the next week and the results totally change. Or the results just do not make any sense and they are clearly contradicted by well-run incrementality experiments, either geo-lift or user-level ones. These tools just clearly are not delivering on the promise that people think that they are making—this idea that we can generate a causal relationship or identify a causal relationship between the marketing activity on one side and the business outcome on the other.
ES: That is really the fundamental issue. It is just people misinterpreting what these tools are supposed to be able to tell them or what they do tell them and assuming it is some kind of replacement for attribution when it is not.
MK: I think that is exactly right. I say this as a vendor in the space because we sell a marketing mix model, but most businesses do not need an MMM. In fact, an MMM will be net negative for them because it will actively mislead them. Last-touch attribution has lots of problems, but for a lot of businesses, it is way better than anything else that they have and an MMM would actually be worse for them than just following the last-touch attribution.
ES: That is a spicy take. We have to let that hang in the air for a second.
MK: You can generate data that shows this. If you imagine that most MMMs are wrong and will actively mislead you because the model is just incorrect and is learning the wrong things from the data, then it is a net negative. Last-touch attribution has a ton of flaws, but especially for smaller businesses that are really performance marketing driven, it works really well.
There have been a couple of studies released recently that showed that last-touch attribution was about 85 percent as good as only doing incrementality experiments. That is an amazing tool to have at your disposal. I have seen this over and over again working with lots of very small businesses personally where it is very clear that they spend money on Facebook and Google and they get a lot more conversions. You can see it just by eyeballing the data and that lines up really well with what you see in last-touch attribution. Last-touch attribution starts to fall down when you have more complex mixes and when you have much larger media budgets, but an average MMM is actually going to be net negative, so it is worse to have the MMM than not to have it.
ES: Do you remember that furor that erupted on LinkedIn a while back about last touch? I had written something about the need for common sense in digital advertising. My point was that if you are spending your first dollar on digital advertising, are you really concerned about the causality of the conversion? Are you really worried about misattributing that? Is that really something that you should be investing time and money into in terms of doing a holdout or some sort of deeper incrementality analysis? If it is not, then we start talking about this second dimension here of just common sense of realistically how likely is it that I am misattributing the effects and my dollar spend is not truly incremental?
That gets more realistic as you scale spend, you have word of mouth, you develop a brand, you develop loyal customers, you start diversifying your mix, and you start having multiple digital channels being run at the same time. But even if you are spending a million a month only on Facebook, do you really need to do some disruptive incrementality study where you just shut off? You might say yes, you do need to do that, and a million is a number that necessitates that. But the broader point is there is that second dimension of just using some common sense here on how likely is it that there are competing claims for these effects. At the first dollar of spend, maybe we would both agree there are none and most people would agree there are none. Then what is the number? What is the dollar number?
MK: Every business needs to sit down and think about this. If you are spending a million dollars a month mostly on Facebook, it is likely your whole business is driven by Facebook. It is probably 100 percent incremental and there is not very much baseline. That sort of thing can push people toward doing some sort of incrementality experiment—shut off or spend up additional dollars in some number of geographies to try to get some signal. How far off are we? Are we off by 5 percent or 10 percent or more than 20 percent? If it is more than 20 percent, let’s dig in and try to figure out how much and what the size of the pie here is. If it is like 5 percent, does it matter? Is it worth spending a bunch more time and energy on this? Probably not.
MMM is even worse because it is so hard to validate if an MMM is correct or not. You can always run an MMM and always get some number out on the other side, but how do you know if it is right? That is the hard problem that is actually very difficult to answer, and especially it is very difficult for a marketer to be able to answer. Is that actually adding any signal? Is that actually helping you make the next decision? In most cases, no. That is why I look at people who are these fairly small companies running these MMMs and they are getting what effectively is just a random number generator out on the other side and they are frustrated by that. Of course it does not make any sense. You have not thought about how to evaluate if the results of this model are even usable or not.
ES: I want to make something clear. I am not claiming that no one should ever run an incrementality test. Incrementality is the gold standard and you should be doing incrementality testing. My point is more about scale and when that becomes mission critical. For some companies, of course, it is mission critical and you have to be doing it. My point is more about the scale question and relating back to when an MMM actually might be net negative in terms of measurement value. That raises a good point. There are these flaws in the open-source approach and the interpretation or the utilization of these open-source models. How should a marketing team evaluate an open-source MMM? What are they looking for? Are they all roughly fungible or interchangeable commodity tools?
MK: The answer is no. The way that you know that the answer is no is you take your data and you run it through two different MMMs, or run it through Meridian with two different settings. You will see that you will get two very different sets of results generally. One run will tell you Facebook is your best channel and you should spend a ton of money on Facebook. Run two will say TikTok is your best channel and Facebook is terrible.
Immediately you know that this is not really reliable. We made one small arbitrary change to the model configuration, or we switched between two different libraries that supposedly do the same thing, and we got totally completely different results. The first thing that any team should be thinking about as they are thinking about whether or not they should use an MMM or which one they should use—and this is true of whether it is an open-source approach or working with some other vendor—is how are we going to evaluate if the model is correct or not.
Everything else does not matter if the model is not correct. If the model is going to mislead us and tell us that a channel is good when it is actually bad and vice-versa, then the shiny reports and the dashboards do not matter. It can only be valuable if we think that the model is actually right. Every brand needs to think about how they are actually going to evaluate that before they start on any MMM project at all.
I think that there are a couple of ways to think about evaluating an MMM. The first I mentioned earlier, which is parameter recovery. We are going to simulate data where we know what is true based on what we believe about how our business works, and we are going to see if the MMM can get the right answer. That is a good way to check if this could even work at all. That is check one.
Check two is being able to predict the future. Can we train this model up to two months ago and then have it predict what is going to happen over the next two months? That is a second way to evaluate if this could plausibly work. That is a better check if you are actually changing your marketing budget over the last two months, because what you want to know is can this model predict what is going to happen as we make changes to our marketing budget. As we spend up on Facebook or spend up on TV or whatever, we want to know that the model can predict what is actually going to happen. That is what it means to have a causal understanding of the data.
The last one is verifying the model with lift tests and experiments. Can we say this model says that TV is really good, and if we go run a lift test, does it actually line up with what the model was saying? Or historically we have a bunch of lift tests; does the model get that right? Those are the ways to actually think about evaluating an MMM, and those are what you need to be doing before you embark on this journey, whether it is open source or whether it is working with a partner.
ES: How should a company approach onboarding a tool?
MK: I would say make the checklist and start running these checks. Start by simulating data. How do you think your business actually works? How do you think marketing actually works in your business? How long are the time shifts? How much does marketing performance change within a channel over time? Simulate the data and do the parameter recovery exercise. If you are onboarding a vendor or working with your internal team, have them do the holdout predictive accuracy check. Only send them data up to two months ago, ask them to forecast the next two months, and then check and see if this is right or if it is terrible. If it is terrible, then you should not trust it for anything else.
Just go through and actually do the checks. Unfortunately, it is a lot of work. This is hard stuff. It requires writing a lot of code and there are not great open-source tools for this, but this is where the actually valuable work lies—in setting up the framework for how we are evaluating it. It is very similar with machine learning. Running XGBoost is really easy. Setting up the framework for how we are validating that there is no leakage between our training and our test set, that when we actually deploy this to production it is going to work the way that we expect, and that we are actually getting the estimated benefit from deploying this model into production—that is where all of the really hard work lies. It is the same thing with an MMM, but even worse because the feedback loop is often much slower than it is with traditional machine learning applications.
ES: What are you seeing as the reliable feedback loop there? Are you seeing companies that can do this with weekly feedback, or does it have to be a month, two months, or a quarter?
MK: You can start to get feedback after a week. The problem with something like an MMM is in general the forecasts that it is making are generally fairly uncertain. Within seven days, it can be very difficult. You have to make very large changes to a marketing program to be able to invalidate a forecast that an MMM is making over the course of a week. What that means is that in general, often you just have to wait longer than that in order to be able to see if these changes that we made are consistent or inconsistent with the forecast that we made a number of days ago.
Over only a seven-day time period, unless there is a huge change in marketing activity over that time and specifically huge changes in marketing activity that we believe only has short-term effects or has most of its effect in the very short term, that is the only way that you can actually invalidate a forecast of an MMM over such a short time period. But if you imagine that you have some awareness-type channel that you believe has fairly long periods of effect—30 days, 60 days, 90 days—it is very difficult to invalidate an MMM’s read over a time period that is much shorter than that.
ES: And then not to mention other effects like spillover, which is another whole other can of worms. But if you are running a fully digital regime where all your ads are digital and you are selling digital stuff—so you are an app and you are not fulfilling rideshare, you’re a game or a fitness app—it could be a shorter feedback loop. You could make those adjustments faster and you could see the reactions faster. You could have a shorter feedback loop than a quarter, probably even shorter than a month.
MK: That is absolutely true, especially for very short-term purchases like games or fitness apps. I work with a whole range of businesses that range from very expensive laundry products that are not bought on that schedule. But if you have an all-digital business with a very cheap first purchase price or even a free first purchase price, then you are going to expect much shorter time shifts and you can get that faster feedback. You do still need fairly large changes in your marketing budget in order to see it in the MMM because of the uncertainty bounds in the estimates around channel performance. These econometric models inherently have a lot of uncertainty. They are not very precise measurement tools.
You might have an estimate that the cost per acquisition for Facebook is between $50 and $75. That is a pretty wide range. You need to see a fairly large movement in Facebook activity to be able to see the result to be able to say if this is right or not. The uncertainty in the parameter estimates makes it so that it is difficult to invalidate just because the uncertainty is so wide.
ES: Of course. You would need a dramatic change, but that is not really that different from using a platform’s own tools. If you reduce a budget by 5 percent and you have these massive confidence intervals, you are not going to get any information there. You would have to cut the budget in half or drop it to zero. But if you are running a fully digital regime, you are in a position to do that. It is not like you have IOs that you send out six months in advance with a TV campaign.
MK: Totally, if you are willing to do it. A lot of organizations are scared to make such large changes in their marketing budget. This is just based on a lot of feeling that Facebook has always been our best channel and we are very nervous to make changes larger than 5 or 10 percent because what if that makes it so that we do not actually hit our target for the quarter or whatever. Theoretically it is possible. In a lot of organizations practically it is very difficult to do to make these big swing changes. Obviously we see it happen. There are some organizations that are very willing to do it, but a lot of organizations, especially as they get bigger and they are at the level where an MMM actually makes sense, you get a lot of very natural conservatism that starts to kick in and people are very unwilling to make those sort of changes because they are very risk-averse.
ES: Why did the MMM space become so crowded? I feel like I see an overwhelming amount of MMM content on LinkedIn. I have no idea when that happened because it seemed to be recent. It seemed like there was just some kind of moment where everyone runs some sort of MMM shop now.
MK: I am obviously part of this, but I think a couple of things are going on. One is technology made it possible to do MMMs much faster than what had traditionally been available. If you imagine that legacy MMM vendors did everything by hand—it was a lot of analysts’ time, it was a lot of putting together decks—we do not actually have to do that. We can run these models automatically in the cloud. We are able to take the work out of the analyst’s hands and put it into the computer’s hands, and so that opened up the ability to do this in a really interesting way.
Second, open-source tools come out. Third, a lot of people started to get worried about digital tracking methods, attribution, and last-touch attribution, and started looking around for other tools that could potentially solve their measurement problem. All of those things come together. There is a big appetite for a new way of doing things. People see the new ability to use MMMs to do it, and so they start talking it up. That is what drove this explosion in MMM vendors. Some of them are MMM-first, and some of them are other attribution tools that tack on an MMM because they can just run Meridian in the background alongside their MTA solution and they can say that they have an MMM as well.
The big problem in the industry is that there is no good way for marketers to evaluate the MMM’s quality in terms of what matters. What actually matters is: is the MMM right? Is this MTA tool that is running Meridian or Robyn or their own home-rolled thing behind the scenes generating actual useful output, or is it just a random number generator? Because it is very difficult or impossible for marketers to evaluate that, then they do not evaluate these MMM vendors on what actually matters. They evaluate them on a bunch of shiny dashboard-type features that are easy to see and easy to evaluate but are not actually the thing that matters for actually using the tool to drive profit into the future. That is a big problem that we have in the industry right now. It is really easy to run an MMM, and it is very difficult for buyers to evaluate the different MMM vendors on the axis that actually matters, which is whether it is actually learning true causal relationships in the data or if it is just spitting out random numbers in a fancy dashboard.
ES: That is the challenge here with MMMs. You have to deliberately instrument the experiments. If you do not do that, it is not going to work. An MMM is going to operate across this historical dataset, but it is not prescriptive in that way. It was just weird to me when people complain about probabilistic methods. How do I know what my ROAS is? You don’t. That is the point. You have to test that. The MMM is just a measurement tool. You have to create the experiment that gives you the two things that you compare. That is the whole point. That is a new operating model for a lot of people, and it does not fit with how they function.
MK: That is totally right. It does not fit with how they function, but also, unfortunately, a lot of vendors in this space and hype people on LinkedIn have pitched MMM as a crystal ball. You do not need attribution anymore; just run an MMM and it will tell you the true incremental value of every single channel and campaign you have. That is just not true at all. It totally oversells the capability of this technology and it sets people up for failure. CMOs or VPs of marketing hear that and they are like: great, I will just buy whatever shiny MMM, it is going to tell me the truth and all my problems are going to be solved. That is just totally unrealistic, but for marketing leaders who are not steeped in the history of econometrics and the history of MMM, it is easy to see how they have been misled by snake oil salesmen really dramatically over-claiming what this technology can do.
ES: Selling it as attribution, selling it as forward-facing attribution versus just measurement, which is very different. We last spoke, I forgot to look this up, but it must have been like two and a half years now. How has the measurement landscape changed since we last spoke?
MK: I already alluded to this at the top of the call, but I think we have already seen peak MMM and I think we are now on the downward slope again. A lot of people got really excited about marketing mix modeling, they over-promised what it could do, they brought it into their organization, it failed fairly dramatically, and now they are either writing it off or starting to realize that actually this is not a crystal ball. This is not actually going to solve all of our measurement problems, and they are starting to really rethink what the role is in their organization.
Obviously, the other big change is AI happened in the last three years, which has put a lot of power into a lot of people’s hands to be able to run their own MMMs, set up their own measurement stack, and be able to really interrogate the numbers in a way that really was not possible before. That is really interesting and exciting in that it allows some companies, at least those that are willing to put the work into doing it, to really start to evaluate these tools on the dimensions that matter and be able to see up close and personal that if I run this MMM with two slightly different datasets or two slightly different assumptions, I am getting different results. That means that it cannot be a crystal ball, and we need to think about a different way to use this tool and to use this paradigm.
Maybe the last major trend is the rise of incrementality testing. A couple of new vendors have popped up and are having a lot of success, which I think is great. People are starting to think about where deliberate experimentation fits into the way that we strategically want to operate a marketing program, how much budget are we willing to commit to testing budgets and to learning about the true value of additional investment into these channels. That has been really great and really positive for the industry overall and a much better foundational base to build on.
ES: Talk to me about the AI analysis piece because that is really interesting. I do think that you have got a lot more people empowered with the ability to be like a marketing data scientist than could have done that in 2023. That to me does feel like a material change.
MK: Absolutely. You have all of the problems that the influencers on LinkedIn constantly talk about—there are cases where data analysis can be challenging and confusing and LLMs get it wrong, so you have to pay attention. But broadly, I see huge amounts of people who historically were not even analysts—VPs of marketing, directors of marketing—sitting down and really diving into the data. I think that it is really positive to the extent, and especially this is what I look for and what gets me really excited, that they start to find flaws in their own assumptions.
They can do that much faster than they were able to do in the past. In the past, you might imagine that for a lot of VPs of marketing, the data that they get comes through in PowerPoint slides, they see it once every couple of weeks, they are not ever really getting into the weeds. They are not ever really forced to confront the potential misconceptions that they might have or the inconsistencies in their own beliefs. As they are able to actually get into the data, they can see that much more clearly. They are able to put the pieces together to say: wait, if Facebook is actually that performant and we just increased spend on Facebook by 30 percent, why didn’t our sales go up? They can start to really dig in and dive into those types of questions, which I think historically they were maybe vaguely aware of but did not have the time or the ability to actually ever dive in and check that out. Now they are able to see it face-to-face with the data via Claude or Codex or whatever. That is a really powerful new thing that is happening that is getting people to ask themselves and their organization more of the really hard questions that they need to be asking.
ES: There are obviously risks here. You get somebody diving into the data who does not have the fluency with even basic statistics, and it could go awry. But I think in general it is a very good thing, and I have seen that across the board. In marketing, there is a very acute application, which is just can I understand the incremental impact? The fact that you can interrogate that question very easily with a couple of prompts, maybe seems like for some marketing teams where it did not exist before, there is more of this nexus of scrutiny, which is nice. That probably at least nudges people away from last click where they might have been totally reliant on it before.
MK: I have seen things even where marketers are like: oh, we have not been able to get our data to run an MMM and that was the big blocker. And then they are like: I just asked Claude to do it, and Claude pulled all of the data. Now it is there, it is done. It went from being a thing that felt totally impossible to them to 30 minutes later it is done. It is hard to understate how much that opens up the ability to get to the next interesting strategic step, and I think that we are only starting to see the impacts of that, but I think it is going to be huge.
ES: How are advertisers adapting their MMMs to CTV? I feel like that is the big growth channel for a lot of direct response marketers. They are seeing CTV as this scale channel and that necessitates the use of an MMM for the first time, or it necessitates that you onboard a new channel and the numbers are going to change, but they are going to change more slowly. You have got to be prepared to invest while maybe the MMM adapts over the course of a month or so. How are they adapting to CTV while using an MMM?
MK: One of the very nice things about MMMs, at least in theory, is that they are channel agnostic. If you can get the data about the channel activity, whether that is spend or impressions by day, you can just slide that into an MMM and it will sort of just work. Whether it is CTV or some other future new channel that has not even been invented yet, you can always just slide it into the MMM and everything will sort of just work. That is a very nice feature of MMM.
You are totally right that as companies are moving into doing more CTV measurement, that often spurs them to consider if we should be using an MMM. The answer is maybe, depending on the business, depending on how big the investment is. For most businesses, one of the nice things about CTV, connected TV in particular, versus something like traditional linear television, is that it is possible to run geographic level experiments with it. That is what I would always recommend doing before starting to do an MMM.
Linear TV is very difficult to experiment with because of the way that it is bought. National TV has a very different buying mechanism than local TV, so it is very unclear even if a geographic test on linear TV bought locally extrapolates to national. But CTV does not. You are able to naturally target at the geographic level, and so that makes it very easy to test and learn into in a way that is going to give you valid incrementality results. What I recommend to most businesses is first just launch it and see if there is any signal at all. Look at your vendor’s measurement panel, always take it with a grain of salt, but take a look at it. Look at your post-checkout survey results. Is anybody saying that they have heard about you from TV? If there is a sign of life there, start to think about scaling it up, and then start to think about how we measure the incremental impact via something like a geographic lift test.
Only at the point where we cannot run tests fast enough to be able to continuously measure this channel, or we need to be able to incorporate all of our different results into something that we can use for forecasting and planning into the future—only at that point should you really be thinking about doing an MMM and using it for all of these different things that are not already solved by our touch-based attribution system, post-checkout survey, and some amount of incrementality testing.
ES: The nice thing about CTV is you can stop and start it. You can scale it up.
MK: It’s great. You can do it on a dime, you can do it in different geographies. It is so much easier to work with and to get started with than traditional linear television.
ES: Is it harder to parse out the effects because of the second screen phenomenon? You can really only do measurement here with a probabilistic model. You might just end up over-attributing to Meta because if someone is seeing the ad and then they are scrolling Meta and they see another ad, that is it. That is what it looks like drove the conversion. How are you seeing the marketers be able to parse that apart?
MK: It is a challenge and the bad news is that it is not really solvable to the degree that people would like it to be. There will be some amount of leakage with an MMM or with an incrementality experiment. You are going to have really wide uncertainty bounds around what your estimate for the impact of that is, and you are never going to be able to get it very precise. The answer is just you are just never going to know.
What you have to do is you have to cobble together signals that are all going to be imperfect as best you can and then try to use that to make a decision. You are never going to know if the true cost per acquisition from CTV is $64 or $68. You are going to get ranges of well, it is probably between $50 and $75, but our post-checkout survey responses around TV have been trending up, so we are going to have to just jump and say we hope that it is actually on the low range of that CPA so it is more profitable for us. We are going to keep doubling down into it because strategically we believe it is important and we want to diversify away from Facebook. Those are the sorts of decisions that you just practically have to make. You are never going to get the fine-grained precision around the estimates of the performance that most marketers would like and that would make the decision really easy.
ES: I would argue that they were overly confident in those metrics beforehand anyway. That ROAS number was never some discrete number, and that is part of the issue with the last-click attribution tools in particular. They just acclimated customers to seeing a single number and having a false sense of certainty around that. That was never justified.
MK: That is totally true. A lot of people at least now intellectually recognize it, even if they do not like it in their heart of hearts. What I point people back to is that you should be thinking about doing forecasting. If you are using these numbers—the last-touch numbers—and you are plugging them into your forecast and you keep missing your forecast, that means something is wrong and those numbers are wrong. You need to go back to the drawing board and think about making your beliefs consistent.
Your beliefs in terms of marketing performance should be consistent with the top-line results that you see at the end of the month. If they are not consistent today, you need to go back and think about what has to change to make those consistent. If you put your focus there, it starts to alleviate a lot of the problems that I think we have seen in the industry of people overly focused on last-touch attribution when it is not actually delivering the results that people think that it is delivering for them.
ES: Michael, this was great. Thank you so much for coming on on short notice. How can people interact with you? How can they consume your wisdom elsewhere?
MK: Follow me on LinkedIn, Michael Kaminsky. We also have a great YouTube channel that we have been putting out a lot of great content on, the Recast YouTube channel. Check me out on LinkedIn and YouTube.
ES: Michael, appreciate your time.
MK: Eric, thanks so much for having me. Great conversation as always.
Comments: