Fraud detection and data quality are not the same thing. While many researchers and audience providers treat them as synonymous, the terms refer to different questions. Fraud detection asks whether a participant is real and acting in good faith, while data quality asks whether what they told you is accurate.
Tools like IP filtering and device fingerprinting were built to answer the first question, but were never designed to answer the second. So how can researchers answer it?
Going beyond fraud detection means changing what determines who gets tested, and why. The identity signals that trip fraud detection tools aren’t enough to help researchers understand who is engaging with the research instrument, and whether their responses accurately reflect them.
Behavioral segmentation is the next step for online market research. While fraud detection boots out obvious bots and known bad actors, segmentation works as a fine-toothed comb that catches the stragglers. In some cases, that means spotting more bad actors; but in many cases, it means finding genuine participants who ended up in the wrong place.
Together with traditional fraud detection, it allows for greater control over the data that surveys produce, with major implications for survey data quality. Here’s how it works.
The Bad Actors Traditional Fraud Checkers Overlook
It’s assumed that once the fraud checkers have done their job, the answers that research audiences provide are reliable. In practice, this isn’t always the case, which compromises survey data quality. But we’ve found that unreliable responses follow reliable patterns.
At Research For Good, our own participant data showed that people who entered certain zip codes — 10001 for Midtown Manhattan, 90210 for Beverly Hills — didn’t behave like the rest of the population.
On their own, these zip codes weren’t fraud signals. None of these audience members triggered connection or device checks. But their survey responses were inconsistent and suspicious enough for us to determine that they posed a risk to data quality. It begged the question: did these people actually live in these neighborhoods?
There are a couple reasons why someone might report a zip code that isn’t their own. Privacy is one reasonable conclusion; people withhold real details online for all sorts of defensible reasons. But privacy doesn’t explain why Beverly Hills, one of the most recognizable and wealth-coded zip codes in the country, would be the go-to answer for such audience members.
Unreliable responses follow reliable patterns.
In some cases, audience members might be providing aspirational answers: describing the life they’d like to have, not the one they do have. This is a real psychographic pattern we’ve observed. But these over-claimers make up a small minority.
In most cases, we’ve found that these are fraudulent actors inputting attributes that get them into surveys they aren’t qualified to join and collecting incentives across multiple studies.
We dubbed these individuals “operators”: a variety of fraudsters who can slip past traditional fraud detection because they often aren’t using any of the technological deception methods that fraud checkers look for. They’re simply lying the old-fashioned way. Sometimes they’re part of an enterprise, and other times they fly solo, but the manual misrepresentation is the same.
Without an additional detection layer, operators slip through with ease. In low-incidence or specialty audiences, unreliable responses can reshape findings entirely, which makes operators a serious threat to survey data quality.
Behavioral segmentation is a progression from traditional fraud detection. It allows us to recognize, flag, and test for suspicious behaviors.
Using behavioral segmentation, we’ve identified a cluster of behaviors consistent enough across studies that we can now anticipate when these bad actors tend to show up, test them accordingly, and, should their responses line up with what we’ve identified as a fraudulent behavior pattern, flag them as a data risk. Here’s what it looks like.
How We Spot Operators, And How They Affect Your Data
An input like a high-risk zip code is not a red flag on its own. But when certain behaviors follow, we know we’ve found an operator.
Over-claiming is the clearest example. Across our data, we’ve observed that certain respondents in the 90210 and 10001 zip codes claim ownership of luxury vehicles and high-value financial assets at rates that significantly outrun real-world ownership statistics.
A look at auto ownership claims among US males aged 18–65, all of whom passed standard fraud detection, showed stark differences between respondents who passed our behavioral tests and those who failed them. BMW ownership was claimed at nearly triple the rate among test-fail respondents. Alfa Romeo ownership was claimed at more than ten times its real-world rate.
Financial asset data showed even more extreme variation. The share of test-fail respondents claiming no financial assets was a quarter of the rate seen among test-pass respondents. Most asset types appeared at multiples of their realistic prevalence.
This specific behavior pattern is reliable enough that we test for operators within this zip code group using survey questions specifically designed to elicit it. When a participant’s responses line up with the behavior patterns of an operator, we know that we can exclude that participant’s responses from the data.
The key difference in Research For Good’s approach is who gets tested and why.
How Behavioral Segmentation Catches What Fraud Detection Misses
Misrepresentation is not demographic in origin, but psychographic. You cannot identify it through methods like IP filtering or device fingerprinting. Catching that misrepresentation requires different tools — ones that don’t test everyone the same way, like traditional fraud checkers, but instead apply scrutiny where the risk is highest.
Research For Good’s system is called the Quality Interception Point, or QuIP. Rather than applying the same battery of checks to every respondent, QuIP uses behavioral and profile markers to identify respondents who are more likely to exhibit suspicious response patterns before the study begins. For example: a 90210 area code.
Once identified, these audience members receive the attention questions and speed traps that, in typical survey design, are applied to every respondent. But we also serve them questions that are deliberately designed to elicit the kind of aspirational or inconsistent responding that signals unreliability.
Critically, the tests are designed not to be obviously identifiable as tests. Answer lists are carefully managed; anchor points that give respondents somewhere predictable to land are removed; questions mix real and unrealistic scenarios, so there is no intuitive way to determine what the “correct” response is supposed to be.
Respondents prone to over-response, such as operators, will reveal themselves under these conditions by failing the test outright. But the system is not singularly focused on booting the operators. When a participant’s responses show some, but not all, of the patterns of an operator, that indicates that the participant might simply be mismatched to the survey.
With QuIP, we can flag the operators who pose a high risk to the data, wave through the low-risk participants who pass every behavioral test, and make an informed decision about what to do with the participants who fall somewhere in between: for instance, discarding their responses in this study, but matching them to a more appropriate one in the future.
How Behavioral Segmentation Empowers Researchers
When Research For Good applied behavioral segmentation to our own work, our internal rejection rate fell by roughly 40%. Our overall rejection rate — the share of completes discarded for fraud, duplication, or quality issues — sits at approximately 4%, compared to an industry average of around 18%.
The most exciting gain is control. When a study returns results that don’t look right, data interpretation is driven not by a hunch, but by clear definitions for likely-unreliable respondent groups. For example: a diagnostic quota can be run to understand what’s actually driving the anomaly. On a tracker, a small oversample at baseline allows calibration across later waves, so a shift in who answered doesn’t get mistaken for a shift in the market.
None of it requires auditing a file record by record, question by question. The behavioral layer surfaces what the traditional fraud detection layer can’t see, painting the most accurate data picture possible and surfacing viable options.
The goal of behavioral segmentation in market research is not merely to remove more respondents and make up for a matching problem upstream. The goal is to understand more about the respondents who are there.
What’s more, the benefits of this approach don’t end with the study that’s in front of you now. This approach can also help researchers recruit the best possible audience for their next study.
For an idea of how respondent matching improves survey respondent quality, read our next article: The Right Respondent For The Right Study.