Data quality is not a filter, it's a discipline with Molly and Stephanie

Description

Every research vendor claims their data is clean, but few can actually measure it. Stephanie Vance and Molly Strawn-Carreño have a host-only conversation about the real business issues behind data quality: the costs of bad data, the difference between a filter-first and discipline-first approach, and how researchers can evaluate and improve the integrity of their work.

Stephanie and Molly walk through a four-layer framework to evaluate and improve data integrity. This model covers everything from survey instrument design to the final proof layer where performance is measured and documented. They also address why standard industry filters are missing a significant portion of fraud and what modern threats like AI-generated responses and response farms mean for the field.

By focusing on design and transparency, insights professionals can better defend their numbers and earn a permanent seat at the decision-making table. This conversation is a call for the industry to stop relying on claims and start producing auditable proof.

Episode Resources

Transcript

Molly - 00:00:01:  

The first player that we're talking about here is prevent, which is about the survey instrument itself. And the question to ask is, is the structure of the survey working for me, or is it working against me? So, for example, a 20-minute survey on a phone is just a fatigue factory, and fatigue is actually the largest single source of unattentive responses, which is bigger than fraud. These are respondents who are genuinely trying to give their feedback. You've just made it a little bit too difficult for them to do so. So, looking at that structure and making sure that you've set up an environment for well-meaning respondents to find success.

Molly -  00:00:42: 

Hello, fellow insight seekers. I'm your host, Molly, and welcome to The Curiosity Current. We're so glad to have you here.

Stephanie - 00:00:50:  

And I'm your host, Stephanie. We're here to dive into the fast-moving waters of market research, where curiosity isn't just encouraged, it's essential.

Molly - 00:00:59:  

Each episode, we'll explore what's shaping the world of consumer behavior from fresh trends and new tech to the stories behind the data.

Stephanie - 00:01:07: 

From bold innovations to the human quirks that move markets, we'll explore how curiosity fuels smarter research and sharper insights. 

Molly - 00:01:16: 

So, whether you're deep into the data or just here for the fun of discovery, grab your life vest and join us as we ride the curiosity current. 

Stephanie - 00:01:28: 

Data quality has been a claim for a long time. Every vendor says their data is clean. Every platform says its respondents are real. But what if you could actually measure that? Today, we're talking about what it looks like when an industry stops asserting quality and starts proving it. 

Molly - 00:01:46:  

So, today, no guest. It's just Stephanie and me. We've been spending a lot of time lately thinking about data quality, not just as a feature that vendors can check off the list, but as a real business issue, one that shapes decisions, timelines, stakeholder trust, and whether insights teams are gonna get a seat at the table. We wanted to talk through a framework that's been shaping how we at aytm think about this.

Stephanie - 00:02:08:  

Yeah. And there's a bigger shift happening right now that I think is really worth unpacking. This conversation, honestly, has been running on claims for years, and that's starting to change. And so I'm genuinely excited for us to get into it today, Molly.

Molly - 00:02:23:  

Me too. And I'm just gonna start right off with saying something that I feel researchers know intuitively, but not necessarily something that is said out loud or comes out in every conversation, which is that the cost of bad data isn't just the cleaning bill to fix it. That's actually probably the smallest part of that.

Stephanie - 00:02:42:  

Absolutely. I love that. Let's start there, Molly. Unpack what, like, some of those costs are for us.

Molly - 00:02:48:  

Yeah. I feel that there's four costs that doesn't exactly show up on an invoice per se, but is still there. It is weighted as a cost, which is, first of all, wrong decisions at scale, reruns and rework that you have to pay twice the money for, and it often takes up to three times as long to get that done. Stakeholders' trust is not cheap to lose, and it's very expensive to rebuild and decision velocity. So, when teams don't trust the data, they ask for more, and that whole organization slows down. And I'm just gonna sort of leave it at that because I have some thoughts on this. But what do you think, Stephanie, is from your experience in the seat with clients and in the seat with researchers, what is the most impactful and the costliest in your experience, and in practice?

Stephanie - 00:03:38:  

Sure. And, I mean, I think first of all, reasonable people can agree with that. I wanna start by saying that it's easy and tempting, for me anyway, to wanna say wrong decisions because, like, that's at the heart of what we're doing. But for me, in reality, it's the erosion of trust, whether that's between the supplier and the client, which is where I have historically felt it as a supplier-side researcher, or between that brand-side researcher and their internal stakeholders. 

Molly - 00:04:08:  

Okay. 

Stephanie - 00:04:09: 

Beyond that, I think there's also this compounding effect that we should talk about. Each of these costs makes the others worse. Wrong decisions erode stakeholder trust. Eroded trust slows decision velocity. Slower velocity creates pressure to run more studies. That creates more opportunities for the same problems, and we're in this data quality loop of hell.

Molly - 00:04:32:  

Yeah. Circle of hell. And that behavior change and how it compounds off of one another and how it changes the way that the workflows work at the organization level, not just the individual users' level, is what actually is the issue, and never shows up in a postmortem, has never shown up as this is the core problem that's fueling and spinning off all these other issues.

Stephanie - 00:04:58:  

Absolutely. That postmortem is usually this very Band-Aid solution of a few people showing up to the table, a couple getting their hands slapped, and then a change in process, and then we're moving on from there. But in reality, to your point, it is a cost that everyone pays. The insights orgs lose trust. The brand manager loses confidence, and the business loses that decision velocity. So, it's critical. 

Molly - 00:05:24:  

Yeah. It's critical work. 

Stephanie - 00:05:25:  

There is a phrase that really resonates with me in aytm's data quality philosophy, and it goes like this: “Data quality is not a filter, it's a discipline.” Those are not the same thing, and I think most of the industry has been optimizing for the wrong one. Molly, to get us started through kind of thinking through this, walk us through what filter-first thinking looks like.

Molly - 00:05:48:  

Yeah. So, this filter-first thinking is really interesting because a lot of times you think we're just gonna let it happen, we're just gonna accept the contamination, and then we'll just remove it on the back end. But it shapes your story along the way. It can't be something that happens at the end. It needs to be part of a whole process. But when 70% of fraud overall that we're seeing as an industry standard slips through just the normal screening criteria, that's what the research on research data shows from our industry, then the filter is only doing a fraction of the job that you actually think it is. So, the practical consequence of that is that your cleaned end size is not actually the number that you think it is. It's the number that the report has already filtered, sure, but the filter missed a huge amount of what should have been caught. 

Stephanie - 00:06:38:  

Yeah. We chalk it up as noise, and I love the phrase that you said at the beginning, which is, accept contamination. And I think that that is just reflective of the fact that we think we have to just accept this. 

Molly - 00:06:50: 

It's just like the industry norm. 

Stephanie - 00:06:51: 

Yeah.

Molly - 00:06:52:  

It's just the cost to do in business. 

Stephanie - 00:06:55:  

Exactly. Yeah. Whereas, conversely, I think discipline first means building the conditions where good data is the norm and where it can thrive, starting before the survey ever goes live, before any respondents see a question. The cost to run either of these approaches, and this is the part that's interesting to me, is really the same. It's just that the outcomes are wildly different.

Molly - 00:07:18:  

Yeah. So, filter-first thinking is, I feel like a lot of what the industry benchmark is that treats data quality as something that has to happen up front, but is not something that is part of the full process. It's sort of just accepting that contamination, right? It's accepting that bad data exists, and it is what it is, and then just try to remove it on the back end. But the issues that we're finding is a lot of the research on research from our industry is showing that 70% of fraud actually slips through those standard methods of cleaning, and that's always changing, and fraudsters are always getting savvier and more innovative. So, the job that you're doing at the front is actually just a small fraction of what it actually should be. It's a completely different mindset. The practical consequence of just thinking with the filter-first mindset is that your cleaned end size is not necessarily exactly what you think it is. The number on the report has already been filtered. But, again, like I said, that filter has missed most of what should have been caught and still creates issues. So, that's a great point, Stephanie, about how it actually needs to be more of a holistic process, but what does it actually look like? Again, I'm gonna ask you in practice questions because I feel like that's what I would wanna know, super valuable as a practitioner. What does it look like when a research team has actually internalized the discipline-first approach, and what's the difference perhaps in how they set up studies? How do they choose vendors? How do they present results? All those kinds of good things.

Stephanie - 00:08:47:  

Absolutely. Well, I think the main thing about this shift is that it changes where in the research life cycle, you invest attention and action related to data quality. Filter first, to your point, puts weight on post-collection processes. Discipline first puts weight on design; it puts it on panel selection, instrument quality, earlier in the process. That's a shift that I'm just firmly convinced we need to see across the industry. And I would turn a question back to you, Molly. What do you think it takes for the industry to move toward discipline first as the default? Is it a buyer change? Is it a vendor behavior change? Is it both?

Molly - 00:09:30:  

I think the answer is actually very simple, which is, it’s just what does the market asks for. It's the market dynamics. It's the economics behind it. Clients really have to demand this and be making vendor selections, making choices about what types of research they're using and not using, and in what ways they have to demand this in order for vendors to truly feel the pressure to respond. There are absolutely vendors out there that are doing this because it's the best way to present data. Ethically, we have to present the best data that we can, and so we're going to adopt this mindset, but there does need to be that external push to manage change in order to say, “We have to make this investment. We have to make this change. It can't be business as usual because we're feeling that pressure to change.”

Stephanie - 00:10:16:  

Yeah. And I'll just add to that. You know, I have never met a supplier-side researcher who didn't care about data quality. That's not what's going on. I think where that could go out the window, though, is when it's made into a competing priority with cost or speed. So, I agree with you that a change in demand from the brand side is key because it allows that priority to surface as something that we get to, that we need to pay attention to.

Molly - 00:10:44:  

Yeah. I mean, the life of a researcher is hard. You have so many competing things that you're balancing, and this type of mindset takes time, it takes effort, and it takes workflow change, and then you have to communicate that change, right, to your clients and your stakeholders. This is gonna take more time, or this is going to be a different approach to this, and having to manage that conversation is also not easy.

Stephanie - 00:11:05:  

Right. Let's revisit this instrument. It's something when a client is done and ready to hand it off for programming; that's not always a conversation they're ready to have at that point, but my goodness, better to have it than two weeks later when we're analyzing data, and it doesn't make a lot of sense to us.

Molly - 00:11:22:  

Yeah. Yeah. So, I wanna talk a bit more about the framework that we've talked about a bunch. And what I find useful about the framework that we've been working with here at aytm is that it gives a specific place in the process where you can look to if something goes wrong or there's a challenge that you're facing, and when you're evaluating whether it's something that will go wrong. So, the four layers are prevent, protect, purify, and prove. And we heard this a little bit because our head of data quality here at aytm, Jonathan Goodbread, recently spoke at the 2026 Quirks Virtual Session – Ensuring Data Quality, Security, and Ethics, which was discussions around data best practices for better research. He introduced this idea about the 4P and the 4P framework, which I think is really helpful for vendor evaluation as a whole. So, let's just dive right into it. The first layer that we're talking about here is Prevent, which is about the survey instrument itself. And the question to ask is, is the structure of the survey working for me, or is it working against me? So, for example, a 20-minute survey on a phone is just a fatigue factory, and fatigue is actually the largest single source of unattentive responses, which is bigger than fraud. 

Stephanie - 00:12:44: 

Yeah. 

Molly - 00:12:45: 

You know, these are respondents who are genuinely trying to give their feedback. You've just made it a little bit too difficult for them to do so. So, looking at that structure and making sure that you've set up an environment for well-meaning respondents to find success.

Stephanie - 00:12:59:  

Yeah.

Molly - 00:13:00:  

The second one, Protect, is about the sample. The question you should be asking is, is everyone who's in this actually really who they say that they are? Because every fraudulent respondent, you can actually think of it actually cost twice, right? It costs once to pay for when you were thinking that they were actually there and well-meaning, and they were gonna give you some good data, and the second cost is to clean the data when you take them out. So, the goal is not necessarily to block them and remove them after the study because you've already incurred the cost of having them in your study. The point is to protect them from entering your study at all, preventing them from coming into that study at all.

Stephanie - 00:13:40:  

Absolutely. And then to jump into that next P, Purify, to me, is where 2026 is genuinely different from even two years ago. The threats have evolved. AI-generated open ends, response farms, which are real humans coordinated sharing infrastructure, and the attention ceiling that you were kind of referencing, Molly, that good design can reduce, but it cannot eliminate.

Molly - 00:14:04:  

Mhmm.

Stephanie - 00:14:05:  

And if a vendor's infield strategy only addresses one of these three, it's solving at best a third of the problem.

Molly - 00:14:12:  

So, it's not gonna work.

Stephanie - 00:14:14:  

It's not. And then finally, the Prove layer is philosophically the most important to me. It's the difference between our quality is good, which is a claim, and something concrete that a CMO can actually read, compare, benchmark, and hold you accountable to.

Molly - 00:14:33:  

So, what's interesting is that you said that Prove is the most important layer philosophically in the process. However, it's the one that can get skipped the most often, which is really interesting to me because it's the one that, you know, is gonna be the most defensible. It's the one that matters when a stakeholder like your CMO challenges your numbers or challenges any of the takeaways here. Without a per-study quality record, you can't defend anything. You have no leg to stand on when they start asking those very specific questions. So, without that comparability across studies, you can't tell if it's your program is getting better, or if it's just getting bigger. You're just doing more things.

Stephanie - 00:15:13:  

And, you know, one of the things that these 4Ps holistically do is that they allow you to sort of assess what we call a data quality report at the study level for every study. And it raises the question for me, what would it change? What would it look like for an insights team if they assessed data quality? They read that data quality report for every study before they even looked at the results, every single time. And that is the action that, yeah, John, our head of data quality at aytm, kind of closes his talk around. I think it's harder than it sounds, but I think this is exactly where, as an industry, we need to be heading, as a discipline we need to be cultivating.

Molly - 00:15:57:  

Well, let's get specific here. 

Stephanie - 00:15:58:

Yeah. 

Molly - 00:15:59:  

Let's just say I am a senior insights director, and I'm sitting across from a research vendor, or evaluating one. I'm in the evaluation process. What are the questions that I should be asking when it comes to data quality? And maybe even more important, what are the types of answers that I should just refuse? Like, what's a red flag I should be on the lookout for?

Stephanie - 00:16:19:  

Okay. Well, first off, I think we have to be asking suppliers whether their fraud detection is single signal or relationship-based. Single signal catches the obvious stuff, but relation-based detection catches coordinated behavior and that's gonna hit your response farms, things like device sharing, proxy patterns. If they can only describe one signal, your vendor, I mean, they're not seeing the full picture as we've talked about. Also, you gotta ask about what the in-field strategy looks like across AI-generated responses, across response farms, and across inattentive respondents distinctly. If they only talk about one of those categories, again, they are solving for about one-third of the problem.

Molly - 00:17:03:  

And these questions and what we're getting at here is there's a broader signal, and this is a moment that's worth naming. For example, aytm is moving towards publishing data quality figures publicly about all of our different surveys that we have running, and it's the first time that anyone in the industry has put out the actual auditable numbers out in the open. And it changes the conversation from just our quality is good to, actually, here are the numbers and here is the methodology and here is how you can better audit us. And that also begs the question of what happens to the industry then if this transparency completely becomes the norm? Does it raise the floor for everyone, or perhaps it maybe only benefits the vendors who are already investing in data quality? All the business will just go to them if everybody starts talking about this, or will it challenge everybody to step up?

Stephanie - 00:18:00:  

Yeah. I think that's a great question, and I will be the first person to say, I don't have a crystal ball.

Molly - 00:18:06:  

Why not? Nobody?

Stephanie - 00:18:09:  

Yet. I don’t. But the answer to refuse as an evaluator of vendors, of panel is ‘trust us’. That's the answer to refuse. Transparency is becoming the buying criterion. Vendors who can show their work will win. Those who can't should be pressed until they can.

Molly - 00:18:28:  

Yeah. And I think my answer to that is also that the credibility gap will widen as people are being more challenged in this. And insights teams that can defend their numbers and be transparent are going to become indispensable, whereas the ones who can't are gonna be asked consistently to rerun a study until they're may be replaced by a different vendor, and that truly is the business case for caring about all of this.

Stephanie - 00:18:54:  

Absolutely.

Molly - 00:18:55:  

So, we've talked a lot about practical applications, but let's talk about perhaps what an insights professional can do tomorrow. So, Stephanie, what is a question that a researcher should ask their vendor this week? A new vendor, or one that they're evaluating or perhaps a vendor that they have an ongoing relationship with. That's maybe something they've never asked before. What is that question?

Stephanie - 00:19:19:  

Okay. I think that the question that I would suggest that brand side researchers ask their vendors or supplier side who use an external panel is to ask the question, walk me through your data quality framework. I want people to forget nitty-gritty details. I don't wanna hear about panel books, cleaning rates, source blends, ask about the framework. If they are not being proactive and driven by a strong point of view, they're scrubbing data, and that's a Band-Aid. And to the whole point of this conversation, that is just not enough anymore.

Molly - 00:19:56:  

Yeah. And I feel like that if it's checking a box, you'll be able to get a feel by their answer in this that they're just checking a box and moving on with their life and sending you what they have, or they're actually driven and passionate and interested in staying ahead of these topics and that they have an internalized way that they go about this process, and it's not just something that they do and then move on. I feel like this answer, I mean, even what they tell you, but also the thematic behind it, their approach to how they answer it is also gonna be telling as well.

Stephanie - 00:20:29:  

Absolutely. Well, thanks so much for joining us today, as always, on The Curiosity Current.

Molly - 00:20:35:  

If this conversation sparked something new for you, we'd love for you to subscribe and leave us a review. It really does help people find our show.

Stephanie - 00:20:43:  

And this conversation about data quality is not over. We will be back with more episodes, with more experts, with deeper points of view, so stay tuned for that.

Molly - 00:20:53:  

And until next time, we encourage you to keep asking the questions that matter, and we'll see you next time on the current.

Outro - 00:21:01:  

The Curiosity Current is brought to you by aytm. To find out how aytm helps brands connect with consumers and bring insights to life, visit aytm.com. And to make sure you never miss an episode, subscribe to The Curiosity Current on Apple, Spotify, YouTube, or wherever you get your podcasts. Thanks for joining us, and we'll see you next time.