A Rising Research Integrity Threat: AI-generated Screener Responses
Kim Torres
July 11, 2024
As UX researchers, we spend a lot of time discussing and exploring ways to incorporate AI into our work. However, we’ve been noticing in projects that potential research candidates seem to be using the same technology to trick us into letting them join a study they should not be part of. Some are simple to spot, such as a bizarre and amusing paragraph about intergalactic space llamas (we are not kidding; someone spent the time to troll us using a well-known panel service). Others are more subtle, where the answers aligned exactly with what we were looking for but felt too well crafted, too on point.
Tips for Identifying Rogue Respondents
As AI capabilities and access evolve, we expect this issue to become more of a problem in the years ahead. The financial incentive to join a study is just too tempting, and AI-generated responses can now let someone with no knowledge of a complex topic appear to be an expert—the exact type of person you are looking for in that project with a recruitment deadline fast approaching. Many recruitment platforms now offer double-screening or video messaging as a paid add-on to weed out fraudulent respondents, which is a great option as well.
Over the past few months, we’ve developed methods and revised our mental models to address what we are seeing. Here are our top recommendations:
Make your open-ended questions more complex We used to rely on reviewing open-ended questions as a great first-pass method for refining our initial respondent pool. Today, however, a once useful inquiry into the complex use of a tool can be easily generated by a respondent using our own question as the AI prompt. In response, we now tend to create questions that require detailed, anecdotal, context-specific answers to help expose potential fraudulent answers.
Be wary of extremely polished answers Sure, someone may take the time to provide an extensive, paragraph-long, well-structured answer with perfect grammar, but we’ve found these comprehensive, overly formal answers that lack any errors should be closely examined. They are just too good. We often flag anything that looks too good and cross reference with other checks. Sometimes as a team, we take the time to discuss and come to a consensus on this flagged group.
Verify authenticity in other curated, reliable platforms If you’re using a platform that links to participant social profiles, and you should be whenever possible, take time to review their background and assess if their answers align with what you see. We’ve found that someone claiming expertise on LinkedIn in the niche area we are desperately looking for sometimes has an online CV that looks radically different than what we would expect.
Follow up with additional questions One round of screeners often doesn’t work for us anymore. As you listen to your gut and identify the responses that feel fishy, leveraging follow-up questions is a great way to feel more confident in making a determination. Asking the same open-ended question slightly differently and comparing the first and second responses should give the insight you need into including or excluding a participant.
When possible, build a trusted, long-term panel Every salesperson knows it’s harder to find a new customer than to expand with the ones they already have. While research often looks for fresh outlooks, there are many times when tapping into a trusted panel makes the most sense. We love using a dedicated panel when possible, as it drastically cuts the time to recruit for a study, and it allows us to track behavior and attitudes over time. In the age of AI, small, long-term panels are likely to become more common.
To Do: Frequently Revise Your Techniques
The above recommendations are what we use today. Will we modify them in the months ahead? We already plan to. We do not doubt that as AI services get better at mimicking people’s actual responses, it will be trickier to weed out those trying to slide into a study. Prompts will include not only an answer to our now more complex open-ended questions, but they may also add “and make it seem like I’m in a bit of a hurry and have a spelling error or grammar mistake."
What do you think? Do you have other tips for researchers? Let us know at hello@proximitylab.com.
As UX researchers, we spend a lot of time discussing and exploring ways to incorporate AI into our work. However, we’ve been noticing in projects that potential research candidates seem to be using the same technology to trick us into letting them join a study they should not be part of. Some are simple to spot, such as a bizarre and amusing paragraph about intergalactic space llamas (we are not kidding; someone spent the time to troll us using a well-known panel service). Others are more subtle, where the answers aligned exactly with what we were looking for but felt too well crafted, too on point.
Tips for Identifying Rogue Respondents
As AI capabilities and access evolve, we expect this issue to become more of a problem in the years ahead. The financial incentive to join a study is just too tempting, and AI-generated responses can now let someone with no knowledge of a complex topic appear to be an expert—the exact type of person you are looking for in that project with a recruitment deadline fast approaching. Many recruitment platforms now offer double-screening or video messaging as a paid add-on to weed out fraudulent respondents, which is a great option as well.
Over the past few months, we’ve developed methods and revised our mental models to address what we are seeing. Here are our top recommendations:
Make your open-ended questions more complex We used to rely on reviewing open-ended questions as a great first-pass method for refining our initial respondent pool. Today, however, a once useful inquiry into the complex use of a tool can be easily generated by a respondent using our own question as the AI prompt. In response, we now tend to create questions that require detailed, anecdotal, context-specific answers to help expose potential fraudulent answers.
Be wary of extremely polished answers Sure, someone may take the time to provide an extensive, paragraph-long, well-structured answer with perfect grammar, but we’ve found these comprehensive, overly formal answers that lack any errors should be closely examined. They are just too good. We often flag anything that looks too good and cross reference with other checks. Sometimes as a team, we take the time to discuss and come to a consensus on this flagged group.
Verify authenticity in other curated, reliable platforms If you’re using a platform that links to participant social profiles, and you should be whenever possible, take time to review their background and assess if their answers align with what you see. We’ve found that someone claiming expertise on LinkedIn in the niche area we are desperately looking for sometimes has an online CV that looks radically different than what we would expect.
Follow up with additional questions One round of screeners often doesn’t work for us anymore. As you listen to your gut and identify the responses that feel fishy, leveraging follow-up questions is a great way to feel more confident in making a determination. Asking the same open-ended question slightly differently and comparing the first and second responses should give the insight you need into including or excluding a participant.
When possible, build a trusted, long-term panel Every salesperson knows it’s harder to find a new customer than to expand with the ones they already have. While research often looks for fresh outlooks, there are many times when tapping into a trusted panel makes the most sense. We love using a dedicated panel when possible, as it drastically cuts the time to recruit for a study, and it allows us to track behavior and attitudes over time. In the age of AI, small, long-term panels are likely to become more common.
To Do: Frequently Revise Your Techniques
The above recommendations are what we use today. Will we modify them in the months ahead? We already plan to. We do not doubt that as AI services get better at mimicking people’s actual responses, it will be trickier to weed out those trying to slide into a study. Prompts will include not only an answer to our now more complex open-ended questions, but they may also add “and make it seem like I’m in a bit of a hurry and have a spelling error or grammar mistake."
What do you think? Do you have other tips for researchers? Let us know at hello@proximitylab.com.