Conceptio › Archive › NCBI PubMed Central
NCBI PubMed Centralopen access

Comparing the Convenience, Data Quality, Generalizability, and Outcomes of Student- and MTurk-Generated Data in an Experimental Vignette Rape Perception Study.

St George S et al. · ncbi_pmc
NCBI PubMed Central · Papers · License: Open Access
Open Source ↗Direct PDF ↓
computer-science-education
computer science education

Comparing the Convenience, Data Quality, Generalizability, and Outcomes of Student‐ and MTurk‐Generated Data in an Experimental Vignette Rape Perception Study - PMC Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice Behav Sci Law . 2026 Jan 19;44(2):249–264. doi: 10.1002/bsl.70039 Search in PMC Search in PubMed View in NLM Catalog Add to search Comparing the Convenience, Data Quality, Generalizability, and Outcomes of Student‐ and MTurk‐Generated Data in an Experimental Vignette Rape Perception Study Suzanne St George Suzanne St George 1 School of Criminal Justice and Criminology, University of Arkansas at Little Rock, Little Rock, Arkansas, USA Find articles by Suzanne St George 1, ✉ , Edmond Osei Arhin Edmond Osei Arhin 1 School of Criminal Justice and Criminology, University of Arkansas at Little Rock, Little Rock, Arkansas, USA Find articles by Edmond Osei Arhin 1 Author information Article notes Copyright and License information 1 School of Criminal Justice and Criminology, University of Arkansas at Little Rock, Little Rock, Arkansas, USA * Correspondence: Suzanne St. George, ( [email protected] ) ✉ Corresponding author. Revised 2025 Dec 27; Received 2025 Aug 25; Accepted 2026 Jan 7; Issue date 2026 Mar-Apr. © 2026 The Author(s). Behavioral Sciences & the Law published by John Wiley & Sons Ltd. This is an open access article under the terms of the http://creativecommons.org/licenses/by-nc-nd/4.0/ License, which permits use and distribution in any medium, provided the original work is properly cited, the use is non‐commercial and no modifications or adaptations are made. PMC Copyright notice PMCID: PMC13084973  PMID: 41555659 ABSTRACT Vignette experiments assessing rape perceptions commonly use samples drawn from convenient sources, like university students or online crowdsourcing platforms like Amazon's Mechanical Turk. In the current study we compared the ease of data collection, cost, data quality, demographic characteristics, and experiment conclusions across these two sample sources, which were collected in a vignette experiment assessing mock jurors' perceptions of a hypothetical sexual assault. Results showed it was faster but more expensive to collect MTurk responses compared to student responses. Samples varied across several data quality measures, with students passing more manipulation checks and returning proportionately more usable cases than MTurkers. We also found considerable demographic differences between the samples, as well as rape myth acceptance (RMA) and victim blaming attitudes. Findings indicate that experiments' results and implications depend on sample source; pinpointing the factors that consistently influence rape perceptions, including RMA, will require replicating studies using diverse sample sources. Keywords: convenience sample, experimental vignette design, juries, MTurk, rape perceptions 1. Introduction Since the 1980s, scholars have highlighted the role of rape myths in both the persistence of sexual violence and the criminal‐legal system's feeble responses to it. Rape myths—which are “attitudes and beliefs that are generally false but are widely and persistently held, and that serve to deny and justify male sexual aggression against women” (Lonsway and Fitzgerald 1994 , 134)—have been shown to foster victim blame and influence legal outcomes at every stage, including reporting, arrest, charging, and conviction. Individuals who endorse rape myths tend to blame victims more and perpetrators less (Gravelin et al. 2019 ; Grubb and Turner 2012 ; Suarez and Gadalla 2010 ; van der Bruggen and Grubb 2014 ). Legal actors' decisions are similarly affected by rape myths, with factors like victim‐perpetrator relationship and intoxication influencing arrest and charging decisions (Alderden and Ullman 2012 ; Morabito et al. 2019 ; Spohn and Tellis 2014 ; St. George and Spohn 2018 ). Unsurprisingly, rape myths also influence jurors' assessments of guilt (Dinos et al. 2015 ; Leverick 2020 ). Given their influence, researchers have examined how victim, perpetrator, incident, and observer characteristics affect rape myth acceptance (RMA) and impact legal decisions. Randomized vignette experiments, in particular, have advanced our understanding of the specific attitudes and case characteristics that shape (mock) jurors' perceptions of sexual violence. Meta‐analyses and systematic reviews consistently find that RMA is a strong predictor of hypothetical victim blame (Gravelin et al. 2019 ; Grubb and Turner 2012 ; Persson and Dhingra 2022 ; Suarez and Gadalla 2010 ). Furthermore, studies consistently show that men have higher RMA and assign more blame to victims than women (Dinos et al. 2015 ; Gravelin et al. 2019 ; van der Bruggen and Grubb 2014 ), and that mock jurors blame victims more when victims are male, intoxicated, or familiar with the perpetrator (Grubb and Turner 2012 ; van der Bruggen and Grubb 2014 ). Studies have also tied RMA and victim blaming attitudes to sexism, racism, homophobia, and conservative political views (Aosved and Long 2006 ; Davies et al. 2012 ; Gavin et al., in press ; Suarez and Gadalla 2010 ; Trottier et al., in press ). Despite their contributions, vignette experiments of mock jurors' rape perceptions face limitations in validity and generalizability. Internally valid vignette designs often lack ecological validity; watching live trials or participating in group deliberations likely elicits different cognitive and social processes. While more ecologically rich designs are possible, they are expensive (Ross 2024 ). Moreover, obtaining representative samples remains difficult. No national registry of jury‐eligible adults exists, and local lists are often incomplete or biased (Helm 2024 ; Devine 2012 ). 1 As a result, many studies rely on convenience samples such as undergraduates or MTurk workers. While experienced researchers may understand the tradeoffs between sample cost, convenience, data quality, and generalizability when designing rigorous studies with impactful, translatable findings, less experienced scholars may benefit from clearer guidance. The current study addresses this gap by comparing the cost, ease of data collection, data quality, respondent demographics, and study outcomes in an experimental vignette rape perception study that collected responses from two samples: students and MTurk workers, hitherto referred to as “MTurkers.” Given convenience samples are likely to remain common in rape perception research, understanding their strengths and limitations can help scholars make more informed methodological choices and improve the rigor of research on rape myths and rape perceptions. 1.1. Rape Myths, Case Attrition, and Evaluating Jury Perceptions Researchers agree that rape myths contribute to the high attrition of sexual assault cases in the criminal‐legal system. These stereotypes and false beliefs about rape, rape victims, and rapists (Burt 1980 ) remain common among students (McMahon and Farmer 2011 ), the general population (e.g., MTurkers, PettyJohn et al. 2023 ; Thelan and Meadows 2022 ), and criminal justice practitioners, like police officers (Dewald and Lorenz 2022 ; Gavin et al., in press ; Sleath and Bull 2015 , 2017 ). Furthermore, researchers commonly find that RMA influences perceptions of victims and perpetrators in hypothetical scenarios (Garza and Franklin 2021 ; Venema 2019 ). For example, Garza and Franklin ( 2021 ) found that police with high RMA perceived hypothetical victims more negatively and were less likely to say they would arrest a hypothetical suspect. The RMA among police officers, and the role of rape myths in their decisions, specifically, has been documented across the globe, including Africa (Aborisade et al. 2024 ), Asia (Lee et al. 2012 ; Sharma and Hamilton 2025 ), Europe (Fávero et al. 2022 ; Gekoski et al. 2023 ; Hermolle et al. 2024 ), and North America (Campbell and Fehler‐Cabral 2018Campbell and Fehler‐Cabral 2018 ; Salerno‐Ferraro and Jung 2022 ; St. George 2025 ; St. George et al. 2022 ). There is some evidence that the influence of rape myths on police and prosecutors' decisions in sexual assault cases is related to their concern with jury perceptions and decisions. For example, police may be reluctant to arrest suspects in cases involving intoxicated acquaintances because they believe juries will not convict (Sinclair 2022 ; St. George et al. 2022 ). Their concerns are not unfounded. In a meta‐analysis of mock jurors' decisions, Dinos et al. ( 2015 ) found that RMA influenced jurors' perceptions of hypothetical rape victims and their verdicts. Leverick's ( 2020 ) meta‐analysis that added more recent studies confirmed these findings. However, examining attitudes in the Unites States, Gavin et al. ( in press ) found that police and prosecutors had higher RMA than the general public. Hence, police and prosecutors' concerns about juries may be exaggerated. Explaining jurors' responses to sexual assault, including the effects of RMA and related case characteristics, can guide police and prosecutors' expectations, mitigate unfounded concerns, and facilitate more aggressive prosecution. As access to real juries is limited, researchers assessing jury attitudes and verdicts commonly administer randomized vignette surveys to samples of mock juries drawn from college campuses or the community. What we know about jurors' rape perceptions, including the effects of RMA and case characteristics on their perceptions, depends on the methods employed by researchers. The sample, specifically, may affect study conclusions. Respondents' level of RMA depends on the sample source, as comparisons of criminal justice practitioners to students (Sleath and Bull 2015 ) and community members (Gavin et al., in press ) have shown. Importantly, Sleath and Bull ( 2015 ) found that police officers and students endorsed different rape myths. Furthermore, studies' conclusions, such as which case or victim characteristics affect perceptions, may likewise depend on sample. Unfortunately, while scholars have compared the RMA of criminal justice practitioners to students (Sleath and Bull 2015 ) and the general public (Gavin et al., in press ), they have not, to our knowledge, compared RMA between students and the general public. So, it is unclear whether students and community members differ in RMA or victim‐blaming attitudes, or if the same case characteristics influence both groups' rape perceptions. This is a problem, as until recently, most research on RMA and rape perceptions of mock jurors used convenience samples composed of student (Dinos et al. 2015 ; Leverick 2020 ), who may differ substantially from the jury eligible population. Confidence in and generalizability of findings, therefore, could benefit from direct comparisons between students and community members on RMA, victim blame, and the effects of experimental variables on jurors' rape perceptions. Unfortunately, we found few studies that compared students and community members' rape perceptions. Specifically, Keller and Wiener ( 2011 ) compared verdicts in a hypothetical sexual assault among students and community members who responded to a newspaper ad, finding that community members were less likely to return guilty verdicts. More recently, Coble ( 2022 ) found that community members, in this case MTurkers, had higher RMA and blamed victims more than students. These studies suggest that sample type may influence study outcomes and researchers' conclusions in rape perception studies. Without studies that directly compare samples, researchers must rely on meta‐analyses to reveal how methodological differences may influence study outcomes. For example, Bornstein et al.’s ( 2017 ) meta‐analysis compared the effect of sample type on jury verdicts in hypothetical criminal and civil trials. They found that over half of all respondents across the 53 studies examined were students, which indicates common use of students in jury perception research. They also found that sample type had no effect on the likelihood of returning a guilty verdict, or the recommended sentence. Unfortunately, meta‐analyses focusing on RMA and rape perceptions, specifically, do not describe the effect of sample type on relevant outcomes. For example, Suarez and Gadalla's ( 2010 ) meta‐analysis assessed the correlates of RMA using 37 studies published between 1997 and 2007, 23 of which relied on student samples and 10 of which relied on community samples. While they reported that RMA was related to education—but not age—they did not report the effect of sample type on RMA. Similarly, Dinos et al.’s ( 2015 ) meta‐analysis on the effect of rape myths on jury decisions found that seven of the nine studies in their sample, published between 1984 and 2011, relied on students. Only one study reported no effect of rape myths on jury decisions, and this study relied on a small sample of community members. More recently, Leverick ( 2020 ) conducted a meta‐analysis of 29 studies assessing mock jury RMA and rape perceptions published between 1984 and 2019 across 11 countries. Only five of the studies included community members, and Leverick did not report on the effect of sample type on outcomes. Trottier et al. ( in press ) also conducted a meta‐analysis of 40 studies on the effects of RMA on rape perceptions published between 2010 and 2023, but they did not even report the sample type, noting instead only mean age or age range and gender composition. While these meta‐analyses offer insights about RMA and rape perceptions, they fail to reveal how sample source—students or community members—affect these outcomes. The current study is not a meta‐analysis, and systematically comparing the RMA, victim blame, and related factors across sample types used in prior research is outside the scope of this literature review. Clearly, such work is warranted. The failure to evaluate the effects of sample source on study outcomes, even in a comprehensive meta‐analysis on RMA and rape perceptions indicates the need for more research like the current study that directly compares different sample types. Comparing samples on relevant outcomes, like RMA and victim blame, as well as the factors that affect these outcomes, can facilitate more robust evidence upon which to guide criminal justice practitioners, who, as noted above, are concerned with the perceptions and decisions of jurors. Our study offers this direct comparison, exploring differences between students and community members drawn from Amazon's Mechanical Turk Crowdsourcing platform. Additionally, we examine other methodological considerations, including the representativeness of samples to the US population, the cost and convenience of data collection, and the quality of data produced. 1.2. Methodological Considerations in Rape Perceptions Experiments 1.2.1. Convenience Samples and Generalizability Relying on convenience samples limits the generalizability of findings about jurors' RMA, rape perceptions, and decision‐making because such samples rarely resemble real jurors demographically (Wiener et al. 2011 ). Students, the most commonly used convenience sample in this type of research (Dinos et al. 2015 ; Leverick 2020 ), differ from jury‐eligible adults in age, income, education, and occupation (Buchheit et al. 2019 ; Hanson 2024 ; Paolacci and Chandler 2014 ; Pew Research Center 2016 ; U.S. Census Bureau 2021 ). They also tend to have less diverse life experience or legal knowledge, which may affect decision‐making processes (Behrend et al. 2011 ). For example, according to the story model of jury decision‐making, jurors “engage in an active, constructive, comprehension process in which evidence is organized, elaborated, and interpreted” based on jurors' own knowledge of similar cases and familiarity with story structure (Pennington and Hastie 1991 , 523). Students may possess less relevant knowledge, which could result in systematic differences in interpretation of evidence or verdicts compared to older, more experienced jurors. Nonetheless, student samples could offer insights into particular groups, such as young adults or those without completed degrees. For instance, Lehmann and Smith's ( 2013 ) study of felony juries found jurors were on average White, 41 years old, and had 15 years of education—slightly less than required for a bachelor's degree in the United States—suggesting overlap with student populations. Moreover, young adults and college students face elevated sexual assault risk (Muehlenhard et al. 2017 ; Mumford et al. 2020 ), potentially making their perspectives especially relevant to sexual assault trials. Still, differences between student samples and jury pools limit generalizability (Wiener et al. 2011 ). Given these limitations, researchers may turn to community samples, which better approximate jury‐eligible adults. For example, Amazon's Mechanical Turk Crowdsourcing Platform (MTurk) provides access a large community participant pool distributed geographically across the United States (Amazon Mechanical Turk, n.d. ; Casler et al. 2013 ). MTurkers more closely resemble the U.S. population on gender, age, and race than restricted convenience samples (Berinsky et al. 2012 ; Casler et al. 2013 ; Levay et al. 2016 ). Moss et al. ( 2023 ), for example, found a representative MTurk sample matched the U.S. population on gender, race, and income. Similarly, Weigold and Weigold ( 2022 ) compared students, MTurker nonstudents, and MTurker students, finding that MTurkers were substantially older and more likely to be men than students. Furthermore, unlike students clustered in one location, MTurkers live nationwide—and internationally!—making them more reflective of national jury pools (Levay et al. 2016 ). However, MTurk samples are not fully representative. They skew younger, more liberal, and more educated than the U.S. population (Berinsky et al. 2012 ; Goodman and Paolacci 2017 ; Moss et al. 2023 ). Thus, both student and MTurk samples diverge from the national population in similar ways—being younger, more educated, underemployed, more liberal, and less religious (Buchheit et al. 2019 ; Paolacci and Chandler 2014 ). Given generalizability concerns plague both student and MTurk samples, researchers assessing mock jurors' rape perceptions must weigh tradeoffs in cost, efficiency, and data quality when selecting between them. 1.2.2. Cost and Convenience Student samples offer several benefits to researchers, particularly low cost and convenience. Students are accessible to university‐based researchers who can recruit through their own or colleagues' courses. In some degree programs, research participation is even required (Lattuca and Stark 2009 ). Online education and survey tools like Qualtrics further expand access. Additionally, compensation is typically course credit (Dickert and Grady 1999 ), making student samples especially valuable for unfunded projects. Because exposing participants to trial records or full simulations is expensive (Bornstein and Kleynhans 2018 ), researchers may prefer to invest in enhancing ecological validity—such as filming mock trials—while using students as a cost‐effective sample. Amazon's Mechanical Turk (MTurk) also offers benefits, primarily in convenience. Researchers can access millions of potential participants worldwide (Amazon Mechanical Turk, n.d. ). They also can recruit thousands of responses quickly, often within hours of advertising the study, and they can target workers with specific characteristics (e.g., U.S. workers, Spanish speakers). By contrast, assembling comparably sized student samples may require multiple semesters. Unlike students, however, MTurkers are paid. While researchers set compensation, large studies can become expensive, especially since MTurkers seek quick, simple tasks (Paolacci and Chandler 2014 ). Low pay may slow recruitment, while restricting samples (e.g., by location) adds fees (Levay et al. 2016 ). Costs also vary with study length and complexity (Chandler and Shapiro 2016 ). So, the actual amount of time to collect data may depend on the complexity and duration of the task, compensation, and participant restrictions. Furthermore, paying participants raises ethical concerns: high pay risks coercion, while low pay risks exploitation, particularly for workers dependent on research income (Millum and Garnett 2019 ). If Institutional Review Boards require researchers to pay participants a living wage, even simple MTurk studies could become prohibitively costly. 1.2.3. Data Quality Beyond cost and convenience, researchers must consider data quality, including effort, attention to manipulated vignette details, and missing data. Comparisons of students and MTurkers on these quality metrics reveal mixed findings. Some studies show MTurkers exert more effort, fail fewer attention checks, and produce more complete data. 2 For example, Kees et al. ( 2017 ) found MTurkers wrote longer open‐ended responses, suggesting greater engagement. They also failed fewer attention checks than students (Hauser and Schwarz 2016 ; Kees et al. 2017 ; Toich et al. 2022 ). MTurkers also tend to skip fewer questions; Dong and Peng ( 2013 ) reported lower missing data rates among MTurkers than the 15%–20% typical in psychological and educational studies. Graham ( 2009 ) likewise found students produced more missing data. Peer et al. ( 2014 ) noted high‐reputation MTurkers—defined as workers with 95% or higher approval ratings—generated especially complete data due to experience and motivation to maintain approval ratings. Other researchers, however, identify quality problems in MTurk samples. Reflecting on their use of MTurkers in several studies, Fleischer et al. ( 2015 ) noted they consistently dropped a substantial proportion of data due to inattention or lack of effort; in one study, they dropped 43% of responses due to inattentiveness. MTurkers also complete surveys more quickly than students (Smith et al. 2016 ; Weigold and Weigold 2022 ) and may fail more attention checks (Goodman et al. 2013 ). Aruguete et al. ( 2019 ) found MTurkers failed more validity checks due to speed, whereas students took longer and made fewer careless mistakes, which suggests students spent more time carefully considering their responses to survey questions. By contrast, while Weigold and Weigold ( 2022 ) found that students took longer to complete their survey, they reported that both students and MTurkers tended to do well on attention checks. Importantly, heterogeneity exists in both samples. Peer et al. ( 2014 ) found only 2.6% of high‐reputation MTurkers failed attention checks compared to 33.9% of low‐reputation workers. Similarly, Chandler et al. ( 2014 ) showed student attentiveness varied with experience in repetitive studies. Data quality also may have shifted over time; Chmielewski and Kucker ( 2020 ) reported data quality declines among MTurkers, even high‐reputation ones. As public familiarity with the platform increases, it is possible that the number of people working on MTurk may be increasing while the quality of their participation—in terms of attention and effort—may be decreasing. Overall, whether students or MTurkers provide better data remains debated, leaving open which is more useful for rape perception research, especially mock jury studies. 2. Current Study As shown above, researchers frequently rely on convenience samples to study perceptions of sexual violence, including experimental vignette surveys testing the effects of rape myth‐related factors on mock jury perceptions of victim blame. In the past, recruitment was constrained to college campuses, but online platforms such as Amazon Mechanical Turk now allow easy access to participants. With more options, researchers must decide which source generates the best data within the limits of time, funding, and research goals. Many studies compare student and MTurk samples on demographics and data quality (Graham 2009 ; Hauser and Schwarz 2016 ; Kees et al. 2017 ; Peer et al. 2014 ; Toich et al. 2022 ; Weigold and Weigold 2022 ), yet several questions remain. No studies have qualitatively compared the cost and process of collecting data. Additionally, prior studies comparing samples on data quality measures often focus on the number of respondents failing manipulation or attention checks without reporting the overall amount of missing data or usable cases, which we argue is an important metric for researchers considering which sample source to use. It is also unclear whether students and MTurkers differ in RMA or rape perceptions, or whether distinct case characteristics influence each group. Further, most comparisons focus narrowly on one of these methodological concerns, rather than the multiple considerations shaping methodological preferences. The current study addresses these gaps. We compare two convenience samples—one composed of undergraduate students, and one composed of workers recruited from Amazon's Mechanical Turk Crowdsourcing platform—on several metrics: (1) the cost and convenience of data collection; (2) the quality of the data produced; (3) respondent demographics; and (4) respondent RMA, perceived victim blame, and experimental predictors of victim blame. This study is exploratory, so we made no a priori hypotheses regarding data quality or sample generalizability. As for study outcomes, we expect that students will have lower RMA and victim blame than MTurkers, given their age and education. However, we make no formal hypotheses regarding the effects of vignette characteristics on victim blame. Rather, using qualitative description and quantitative analyses, we highlight similarities and differences between these sample sources, the data they produced, and the findings they yielded. Our goal is to offer guidance to researchers interested in evaluating jury perceptions of rape. 3. Methods 3.1. Transparency and Openness We acknowledge the limited use of OpenAI's ChatGPT (April 2025 and August 2025 versions) in preparing this manuscript. The system was used selectively to improve grammar, phrasing, and clarity at the sentence level for the narrative sections of this text (not the results section). It was not used to find reference material, format citations, or draft, generate, or conceptualize content, or to analyze data. It did not contribute to the intellectual ideas, analyses, or interpretations presented in this paper. 3.2. Data Availability Data are not publicly available, but may be available from Dr. Suzanne St. George, upon request. 3.3. Data Source and Procedure The current study used data from a randomized vignette survey assessing perceptions of a hypothetical sexual encounter. Comparing sample sources was not an original goal of that study; rather, the purpose was to assess the effects of victim race, gender, and sexual orientation on rape perceptions, including how these factors moderate the effects of rape myths on victim and perpetrator blame. The survey used a 4 (victim race) × 5 (victim gender/sexual orientation) × 5 (rape myth condition) between‐subjects design. An a priori power analysis using G*Power revealed that a sample size of 4100 would be needed to detect a small effect ( f = 0.1). A sample of approximately 5000 respondents was selected to achieve the 4100 needed, with room for missing data and incomplete questionnaires. Given the large sample size required, it was decided to recruit from both MTurk and Arizona State University. A mixed sample allowed for the author of the original study to reduce total cost and increase efficiency of respondent recruitment. A study protocol along with relevant materials were submitted to the Institutional Review Board at Arizona State and received approval in Spring 2020 (Coble 2022 ). Data collection began in August 2020 and continued through August 2021. All questionnaires were completed online in Qualtrics. After indicating consent to participate, respondents were randomly assigned to read one of the 100 vignette versions, each of which described a hypothetical date rape followed by a delayed disclosure. 3 Vignettes varied the victim's and perpetrator's race (e.g., both White, both Black, both Latino/a, neither race described) as well as the victim's gender (e.g., “girl,” “guy,” or “transgender girl”) and the perpetrator's gender (e.g., “girl” or “guy”). Sexual orientation was described when the victim/perpetrator were gay/lesbians, and implied by the victim and perpetrator's gender pairings (e.g., “girl–guy,” and “trans gender girl–guy” = heterosexual; “guy–guy” and “girl–girl” = gay/lesbian). See the supplemental materials for examples of vignettes. The frequency of responses to each vignette version was well balanced (26% White victims, 25% Black victims, 25% Latinx victims, 24% no race specified) indicating that participants were effectively randomized to conditions. Respondents then answered manipulation check questions, followed immediately by questions assessing their perceptions of the victim's blame and responsibility. They also responded to the Subtle Rape Myths scale (McMahon and Farmer 2011 ) and demographic questions. 4 The questionnaire ended with questions assessing experiences with sexual violence. Respondents could not backtrack to prior sections. After completing the survey, respondents were taken to a debrief page that explained the goals of the study and provided links to online resources for victims of sexual assault. 3.4. Variables and Measures We compared the student and MTurk samples across several factors. First, we examined the cost and convenience of data collection. Cost was evaluated in U.S. dollars and is the total spent to administer the survey, which included respondent compensation and MTurk administration fees. We also assessed how much time (hours, months) researchers spent collecting data from each sample source, and some of the steps required in this process. Second, we assessed the quality and usability of the data collected. We assessed the proportion of students and MTurkers who finished the survey. Surveys were counted as finished (yes = 1) if the respondent worked all the way through to the end of the survey. Finished surveys could contain missing data if respondents skipped questions. We assessed missing data in two ways: missing data was the sum of missing values on questions and variables used in the final analyses, including individual victim blame items, RMA items, and demographic characteristics (age, race, gender, sexual orientation, education, religiosity, and political orientation); cases were counted as incomplete (yes = 1) if they were missing information on one or more variables. We also assessed respondents' attention throughout the survey, and to the vignette manipulations, specifically, as well as the time to complete the survey and if the response was usable in the final study. After reading the survey, respondents answered two questions asking about the victim's and perpetrator's race. If the respondent read a vignette in which the victim's and perpetrator's race was specified (e.g., White, Black, or Latino/a), and they incorrectly identified the victim's or perpetrator's race, then they were marked as failing the manipulation checks ( manipulation failure , yes = 1), Additionally, there were two questions embedded within the Likert scale questions instructing respondents to select a specific response (e.g., “Mark strongly agree in response to this question.”). Respondents who answered these questions incorrectly, or entered nonsensical responses to open‐ended questions (e.g., responded “1834” to the question “How old are you in years?”) were noted as failing attention checks ( attention failure , yes = 1). Time to complete assessed the number of minutes respondents spent taking the survey. 5 Finally, cases were counted as usable (yes = 1) if the respondent finished the survey, there were no missing data, attention failures, or manipulation failures, and when the respondent spent at least 5 minutes completing the survey. 6 Third, we compared student and MTurker demographic characteristics, including age, gender, race, sexual orientation, education, political orientation, and religious affiliation (see Table 1 for variables, coding scheme, and sample characteristics). Fourth, we compared students' and MTurkers' RMA and the blame they attributed to the victim. We assessed RMA using an adaptation of McMahon and Farmer's ( 2011 ) Subtle Rape Myths scale. The 22 items in the original scale were modified to be gender neutral and respondents indicated their agreement on a 5‐point Likert scale (1 = Strongly disagree–5 = Strongly agree). Victim blame—the primary outcome of interest in the original study—was the average of six items measured on a 5‐point Likert scale (1 = Strongly disagree–5 = Strongly agree) assessing the victim's consent, responsibility, and control in the incident (e.g., “Annie consented to sex with David.”). Higher scores indicated more victim blame for the assault. TABLE 1. Study variables, coding schemes, and sample characteristics. Variable Coding scheme Students ( n = 538) MTurk workers ( n = 2006) Complete sample ( n = 2544) N (%)/mean (SD) N (%)/mean (SD) N (%)/mean (SD) Respondent demographics Age How old are you in years? (fill in) 24.8 (7.2) 38.8 (12.0) 35.84 (12.5) Gender One question asking gender identity; select all that apply Male 159 (29.6%) 1086 (54.1%) 1245 (48.9%) Female 368 (68.4%) 889 (44.3%) 1257 (49.4%) Other gender 11 (2.0%) 31 (1.6%) 42 (1.7%) Heterosexual One question asking sexual orientation; select all that apply 412 (76.6%) 1613 (80.4%) 2025 (79.6%) Race One question asking racial identity; select all that apply White 409 (76.0%) 1601 (79.8%) 2010 (79.0%) Black 43 (8.0%) 221 (11.0%) 264 (10.4%) Asian/Hawaiian 52 (9.7%) 130 (6.5%) 182 (7.2%) Other/mixed 34 (6.3%) 54 (2.7%) 88 (3.5%) Latinx/Hispanic One question asking ethnicity 156 (29.0%) 90 (4.5%) 246 (9.7%) Highest education Ordinal scale assessing highest level of education High school or less 34 (6.3%) a 142 (7.1%) 176 (6.9%) Some college/AA 412 (76.6%) 402 (20.0%) 814 (32.0%) Bachelor's or higher 92 (17.1%) 1462 (72.9%) 1554 (61.1%) Political orientation One question asking political orientation Progressive/Liberal 276 (51.3%) 832 (41.5%) 1108 (43.6%) Moderate 175 (32.5%) 532 (26.5%) 707 (27.8%) Conservative/very conservative 87 (16.2%) 642 (32.0%) 292 (28.7%) Religious Ordinal scale assessing “how intensely do you practices religion”; collapsed into “NA‐not affiliated”/“not religious at all” = reference, any practice = 1 307 (57.1%) 1124 (56.0) 1431 (56.3%) Study outcomes RMA 22 items measured on 5pt‐Likert scale, averaged; alpha = 0.98; range 1–5 1.72 (0.6) 2.74 (1.1) 2.52 (1.1) Victim blame 6 items measured on 5pt‐Likert scale averaged; alpha = 0.92; range: 1–5 2.02 (0.9) 2.99 (1.1) 2.79 (1.1) Vignette manipulations Victim gender Victim described as a “girl,” “guy,” or “transgender girl” Victim sexual orientation Victim described as “gay,” “lesbian,” or no modifier Victim race Victim described as “White,” “Black,” “Latino/a” or no modifier Open in a new tab a Thirty‐Seven students marked their highest education as high‐school or less. If these individuals were freshmen taking the survey in their first semester of college, they may not have counted their education as more than high school. Finally, we assessed if the effects of the experimental variables (victim gender, sexual orientation, and race) on blame attributions varied by sample type. Victim gender (cisgender man = 0, cisgender woman = 1, transgender woman = 2) was manipulated by describing the victim as a “girl,” “transgender girl,” or a “guy.” Sexual orientation (heterosexual = 0, gay/lesbian = 1) was manipulated by describing some victims as “lesbian” or “gay,” while the “heterosexual” condition contained no additional modifier. Finally, victims were described as “White,” “Black,” “Latino/a,” or with no race modifier. 3.5. Analyses and Organization of Findings Below, we describe the samples' similarities and differences. First, we qualitatively described the cost and convenience of collecting survey responses from students and MTurkers. Next, we used independent samples difference of means t ‐tests and difference of proportion z ‐tests to describe differences between the samples on data quality measures (Table 2 ). These tests do not require the groups, or samples, to be equal in size. Then, we limited the samples to “usable” cases and used t ‐tests, z ‐tests, and Chi‐Square tests to compare demographic characteristics, RMA, and victim blame (Table 3 ). Finally, we used one‐way ANOVAs and difference of means two‐tailed t ‐tests to assess the effects of victim gender, sexual orientation, and race on victim blame first among students, then among MTurkers, and finally, among the combined sample (Table 4 ). Throughout the findings section, we include our interpretations. We believe this organization reduces redundancy, improves readability, and allows us to focus the discussion section on the implications and recommendations for researchers conducting experimental rape perception research. TABLE 2. Comparison of student and MTurk samples on data quality measures. Data quality measure Combined sample ( n = 5110) Students ( n = 939) MTurk workers ( n = 4171) Test of difference N (%)/ mean (SD) N (%)/mean (SD) N (%)/mean (SD) Mean difference/Proportion (SE) 95% CI: lower, upper Test statistic p Effect size Finished 3973 (77.8%) 754 (80.3%) 3219 (77.2%) 0.03 (0.01) 0.00, 0.06 2.08 0.038 −0.029 Missing data 5.64 (12.56) 6.68 (14.00) 5.41 (12.212) 1.27 (0.45) 0.38, 2.16 2.81 0.005 0.101 Incomplete 1173 (23.0%) 249 (26.5%) 924 (22.2%) 0.04 (0.02) 0.01, 0.07 2.87 0.004 −0.040 Manipulation failure 767 (15.01%) 65 (6.9%) 702 (16.8%) −0.10 (0.01) −0.12, −0.08 −7.68 0.000 0.107 Attention failure 1551 (30.4%) 307 (32.7%) 1244 (29.8%) 0.03 (0.02) −0.00, 0.06 1.73 0.084 −0.024 Time to complete 12.33 (8.64) 15.74 (9.88) 11.56 (8.14) 4.17 (0.31) 3.57, 4.78 13.63 0.000 0.492 Usable 2544 (49.8%) 538 (57.3%) 2006 (48.1%) 0.09 (0.02) 0.06, 0.13 5.09 0.000 −0.071 Open in a new tab Note: Difference of mean two‐tailed t ‐tests and Cohen's d (effect size) were calculated for the two continuous variables: time to complete and missing data . Difference of proportion two‐tailed z‐tests and Cramer's V (effect size) were calculated for categorical variables finished , incomplete , manipulation failure , attention failure , and usable . While Cramer's V is typically used in conjunction with X 2 tests of independence, we deemed it appropriate here because when both categorical variables assessed in a X 2 have two categories, the X 2 is functionally identical to the z ‐test of difference of proportions. Either test would be appropriate, but we preferred the z ‐test because it is more intuitive, and the display of test information matches the display of t ‐tests. For both Cramer's V and Cohen's d , values of 0.2, 0.5, and 0.8 reflect small, medium, and large effects, respectively. TABLE 3. Comparing students ( N = 538) and MTurkers ( N = 2006) on demographic characteristics, rape myth acceptance, and victim blame. Variable Test type Effect size type Test of difference Mean difference (SE) 95% CI: Lower, upper Test statistic p Effect size Respondent demographics Age Two‐tailed t ‐test Cohen's d −14.05 (0.54) −15.11, −12.99 −26.01 0.000 −1.263 Gender Chi‐square Cramer's V 102.83 0.000 0.201 Heterosexual Two‐tailed z ‐test Cramer's V −0.04 (0.02) −0.08, 0.00 −1.96 0.050 0.039 Race Chi‐square Cramer's V 26.67 0.000 0.102 Latinx/Hispanic Two‐tailed z ‐test Cramer's V 0.25 (0.02) 0.21, 0.28 17.08 0.000 −0.339 Highest education Chi‐square Cramer's V 640.28 0.000 0.502 Political orientation Chi‐square Cramer's V 52.02 0.000 0.143 Religious Two‐tailed z ‐test Cramer's V 0.01 (0.02) −0.03, 0.06 0.43 0.669 −0.008 RMA Two‐tailed t ‐test Cohen's d −1.02 (0.05) −1.11, −0.93 −21.45 0.000 −1.041 Victim blame Two‐tailed t ‐test Cohen's d −0.97 (0.05) −1.07, −0.87 −19.01 0.000 −0.923 Open in a new tab Note: Variable means/SDs and frequencies/percents are displayed in Table 1 so not repeated here. Comparisons tests use “usable” cases. TABLE 4. The effects of experimental variables on victim blame among students and MTurk workers. Mean victim blame by gender condition One‐way ANOVA test for difference of means Sample type N Cis man Cis woman Trans woman F p Eta‐squared 95% CI: Lower, upper Student 538 2.20 (1.0) 1.92 (0.9) 1.94 (0.9) 5.19 0.006 0.019 0.002, 0.046 MTurk 2006 3.10 (1.0) 2.93 (1.1) 2.92 (1.1) 6.51 0.002 0.006 0.001, 0.015 Combined 2544 2.92 (1.1) 2.69 (1.1) 2.73 (1.1) 11.68 0.000 0.009 0.003, 0.017 Mean victim blame by sexual orientation condition Difference of means two‐tailed t ‐tests Heterosexual Gay/lesbian t p Cohen's d 95% CI: Lower, upper Student 538 2.02 (0.9) 2.03 (1.0) −0.19 0.853 −0.016 −0.189, 0.156 MTurk 2006 2.97 (1.1) 3.04 (1.0) −1.54 0.124 −0.071 −0.161, −0.019 Combined 2544 2.77 (1.1) 2.82 (1.1) −1.06 0.290 −0.043 −0.123, 0.036 Mean victim blame by race condition One‐way ANOVA test for difference of means White Black Latino/a Not specified F p Eta‐squared 95% CI: Lower, upper Student 538 1.97 (0.9) 2.08 (1.0) 2.07 (1.0) 1.99 (0.9) 0.46 0.710 0.003 −, 0.012 MTurk 2006 3.06 (1.0) 2.86 (1.1) 2.88 (1.1) 3.12 (1.1) 7.33 0.001 0.011 0.003, 0.020 Combined 2544 2.85 (1.1) 2.72 (1.1) 2.67 (1.1) 2.88 (1.1) 4.84 0.002 0.006 0.001, 0.012 Open in a new tab Note: Tukey HSD pairwise comparisons of perceived victim blame revealed that, both students and MTurkers blamed cis men more than cis women or trans women, and these differences were statistically significant. Tukey HSD tests also revealed that MTurkers blamed White victims and victims with unspecified race more than both Black and Latino/a victims, and these differences were statistically significant. 95% confidence intervals are for the reported effect sizes. Eta‐squared values of 0.01, 0.06, and 0.14 or higher represent small, medium, and large effect sizes, respectively (Brown 2008 ). 4. Findings 4.1. Cost and Convenience Altogether, it took the primary author, Dr. Suzanne St. George, about 1 year and $3035 to collect 5110 survey responses. There were notable differences between the student and MTurk samples across collection procedure, timeline, cost, and number of respondents. The student sample took much longer to collect than the MTurk sample, included fewer respondents, and cost $0. Beginning in September 2020, instructors teaching various online courses were contacted in the Fall 2020, Spring 2021, and Summer 2021 semesters to encourage them to share the survey link with their students. 7 Follow up emails were sent to instructors periodically to encourage them to share the survey with each (new) class (e.g., beginning of each semester/7‐week sessions starting halfway through the semester). The student survey closed in July of 2021. Altogether, it took approximately 10 months and many emails to collect 939 student responses, with about 94 responses per month at $0 per response. 8 By contrast, 4171 MTurk responses were collected over 6 months (February–August 2021) and cost $3035. The study offered $0.50 for participation and restricted participation to MTurkers living in the United States and who had not previously taken the survey (e.g., one response per worker ID). The survey link was advertised to MTurkers 12 times over the 6‐month data collection period. This was done to reduce the monthly out of pocket cost of administering the survey. 9 Each of the 12 batches cost about $250 and yielded about 350 responses within 2 hours of the survey link being posted. So, the MTurk data collection process generated an average of 695 responses per month at $0.73 per response. 10 Clearly, the MTurk sample was more convenient, as Dr. St. George was able to generate a large sample quickly. In fact, the sample could have been collected within a single day if the researcher was not constrained financially. However, the MTurk sample cost substantially more than the student sample. Even though the compensation ($0.50) and administration fees/taxes ($0.23) per respondent were small, collecting over 4000 responses was not cheap. Furthermore, only 48% of responses were usable in the final analysis (as discussed below), meaning the final, total cost per usable response was considerably higher than originally expected ($1.51). Still, the total compensation paid to respondents was small. As more online survey platforms emerge in response to user demand, competition between platforms may require researchers to offer better compensation. 11 As such, it may become increasingly costly to collect data from MTurkers or from other online platforms. 4.2. Data Quality The quality of data produced by students and MTurkers differed in several ways (see Table 2 ). A larger proportion of students completed the survey ( n = 754, 80%) compared to MTurkers ( n = 3219, 77%). Student surveys had more missing data: students skipped an average of 7 questions (SD = 14), while MTurkers skipped an average of 5 questions (SD = 12). Twenty‐seven percent of students returned incomplete surveys, compared to 22% of MTurkers. However, 17% of MTurkers failed the manipulation check question compared to only 7% of students, and similar proportions of students (33%) and MTurkers (30%) failed attention checks. Students also took longer to complete the survey, spending on average 16 min (SD = 10), while MTurkers spent on average 12 min (SD = 8). Differences of proportion z ‐tests and difference of means t ‐tests, displayed in Table 2 , revealed statistically significant differences between students and MTurkers on every data quality measure except attention failures. However, the effect sizes tended to be very small, indicating these differences were not particularly meaningful. Only the effect size of time to complete indicated a medium sized effect of sample type. Furthermore, when accounting for all the data quality measures together, students were more likely to returned usable surveys compared to MTurkers: 57% of student responses were able to be used in the final analysis, compared to just 48% of MTurkers, and this effect was statistically significant ( z = 5.09, p < 0.001), though the effect size was small (Cramer's V = −0.071). Altogether, these analyses indicate that neither students nor MTurkers produce high quality data, though students appear to be better than MTurkers at attending to experimental manipulations, which resulted in returning more usable surveys. Prior studies comparing student‐ and MTurker‐produced data report mixed findings. Some found students took more time (Smith et al. 2016 ; Weigold and Weigold 2022 ) and answered more carefully (Aruguete et al. 2019 ). Others found MTurkers performed better on attention and effort measures (Kees et al. 2017 ; Hauser and Schwarz 2016 ; Toich et al. 2022 ). Still others found no difference in attention across sample type (Weigold and Weigold 2022 ). Our results indicate both groups face notable data quality issues. Students paid more attention to manipulations and spent longer on surveys but skipped more questions, and fewer returned complete surveys than MTurkers. Ultimately, only 57% of student responses were usable. Although this was higher than MTurkers (48% usable), we cannot say that students are necessarily preferrable, especially when considering efficiency and convenience of data collection. 4.3. Respondent Demographics Next, using only usable cases, we examined differences between students' and MTurkers' demographic characteristics (see Table 1 for sample characteristics and Table 3 for tests of difference and effect sizes). Students and MTurkers differed on every characteristic examined except for religious . MTurkers were on average 39 years old (SD = 12) compared to students who were on average 25 years old (SD = 7). This difference was statistically significant and very large ( t [ df = 2542] = −26.01, p < 0.001, d = −1.26). Additionally, there were notable differences in gender identity across sample type: 68% of students identified as women compared to 44% of MTurkers. A slightly smaller proportion of students (77%) identified as heterosexual compared to MTurkers (80%). Students and MTurkers also differed across race and ethnicity. Smaller proportions of students identified as White and Black compared to MTurkers, but larger proportions identified as Asian/Hawaiian and Other/Mixed. Additionally, 29% of students identified as Latinx/Hispanic compared to just 5% of MTurkers. Students and MTurkers also differed across education and political orientation, but not religious. Perhaps unsurprisingly given they were recruited while actively attending college, 77% of students indicated some college or associate’s degree as their highest level of education. By contrast, 73% of MTurkers had a bachelor's degree or higher. Students also were more liberal than MTurkers, with 51% identifying as progressive or liberal, 33% as moderate, and just 16% as conservative or very conservative, compared to MTurkers who were 42% progressive/liberal, 27% moderate, and 32% conservative/very conservative. Similar proportions of students and MTurkers identified as religious. Collectively, these findings show students and MTurkers differed across nearly every demographic measure, which is consistent with prior comparisons (Weigold and Weigold 2022 ). Students were younger, more likely to identify as women and non‐heterosexual, and more racially diverse—including a higher proportion identifying as Latinx/Hispanic—than MTurkers. They were also less educated and more liberal. Differences in age, gender, and political orientation likely reflect characteristics of typical college students (Hanson 2024 ; Pew Research Center 2016 ; U.S. Census Bureau 2021 ), while racial/ethnic differences may stem from the Southwestern location of the university, which has a large Latinx population, though more online education may uncouple these factors over time. Importantly, neither sample was nationally representative, though MTurkers aligned more closely with U.S. demographics on some factors (U.S. Census Bureau 2021 ). The student sample overrepresented women, while MTurk overrepresented men but was closer to the national gender split. Both overrepresented White people; students underrepresented Black people but overrepresented Asian/Hawaiian and Latinx people. The MTurk sample more closely matched U.S. proportions of Black and Asian/Hawaiian people but substantially underrepresented Latinx/Hispanic people. 4.4. Rape Myth Acceptance and Rape Perceptions Finally, we compared students' and MTurkers' rape‐specific attitudes, including RMA and perceived victim blame. We also compared the effects of the manipulated victim characteristics—gender, sexual orientation, and race—on perceived victim blame. We wanted to know if researchers would draw the same conclusions in their rape perception studies, regardless of sample source, or if different samples would yield different findings. Results were mixed. Specifically, MTurkers had higher RMA than students (mean MTurk = 2.74, SD = 1.1 vs. mean Student = 1.72, SD = 0.6). This difference was statistically significant, and the effect was large ( t = −21.45, p < 0.001, d = −1.04). MTurkers also indicated higher victim blame (mean = 2.99, SD = 1.1) than students (mean = 2.02, SD = 0.1), and the difference was statistically significant and large ( t = −19.01, p < 0.001, d = −0.92). Therefore, students and MTurkers differed substantially in their RMA and rape perceptions. We expected that MTurkers would have higher RMA, and by extension, victim blame, than students, because prior researchers have found that people tend to endorse more rape myths and blame victims more when they are men, older, and more politically conservative, characteristics consistent with the MTurkers in this study (Suarez and Gadalla 2010 ; Persson and Dhingra 2022 ). Still, documentation of this difference is important and novel. Researchers relying on students to assess RMA and rape perception may underestimate the strength and pervasiveness of rape myths and victim blame, given these attitudes were much higher among MTurkers. Findings from the experiment, however, are more mixed (see Table 4 ). One‐way ANOVAs testing the difference of mean victim blame across victim gender revealed a statistically significant but small effect in both the student sample ( F [ df = 2] = 5.19, p < 0.01, Eta‐squared = 0.019) and the MTurk sample ( F [ df = 2] = 6.51, p < 0.01, Eta‐squared = 0.006). Victim gender influenced perceived victim blame in both groups, with both students and MTurkers blaming victims more when they were described as cisgender men compared to cisgender women or transgender women. Additionally, two‐tailed difference of means t ‐tests revealed that victim sexual orientation did not influence victim blame among students or MTurkers. Both groups tended to blame heterosexual and gay/lesbian victims the same amount. However, one‐way ANOVAs revealed that victim race only influenced victim blame among MTurkers ( F [ df = 3] = 7.33, p < 0.001, Eta‐squared = 0.011). Although the effect was small, tests indicated that MTurkers blamed victims more when they were described as White or without any specified race, compared to when they were described as Black or Hispanic. Victim race had no effect on students' blame attributions. They attributed equal levels of blame to victims, regardless of their race. Collectively, these findings indicate that conclusions from rape perception experiments may depend on sample source. Students and MTurkers differed substantially in RMA and victim blame, and the effect of victim race on blame appeared only among MTurkers. Thus, causal claims may be sample‐specific: if we had only sampled MTurkers, we would conclude that victim race influences blame; with only students, we would conclude it does not. It is plausible that demographic composition may explain some of these differences, as attitudes toward sexual violence correlate with gender, age, and political views (Suarez and Gadalla 2010 ; Persson and Dhingra 2022 ), all of which varied between samples. We explored this possibility in linear regression models assessing the effects of demographic factors on (1) RMA and (2) victim blame (see Supporting Information S1 : Tables 1–3). Results revealed inconsistent associations across sample types. Age, education, and religiosity predicted RMA only among MTurkers, while sexual orientation and race predicted RMA in both groups but in opposite directions. Similarly, predictors of victim blame varied: sexual orientation, race, and political orientation were significant only among MTurkers, while Latinx identity was significant only among students. 12 These findings problematize generalizability claims, as differences in MTurkers' and students' attitudes cannot be explained by differences in sample demographics. Rather, sample source itself likely shapes outcomes of rape perception experiments, with implications for generalizability to jury decision‐making. Below, we discussion the implications of these findings for researchers conducting experimental vignette studies on rape perceptions. 5. Discussion Researchers studying RMA and rape perceptions, especially among mock jurors, often rely on convenience samples pulled from student populations or Amazon Mechanical Turk (MTurk). While some studies have compared the RMA and rape perceptions of criminal justice practitioners to students (Sleath and Bull 2015 ) or community members (Gavin et al., in press ), how students compare to community members, specifically MTurkers, remains unclear. Furthermore, meta‐analyses and systematic reviews of the RMA and rape perception literature have failed to comment on the characteristics of different samples, differences between samples, or the strengths and weaknesses of research studies relying on different sample types (See Dinos et al. 2015 ; Gravelin et al. 2019 ; Grubb and Turner 2012 ; Hockett et al. 2016 ; Leverick 2020 ; Persson and Dhingra 2022 ; Suarez and Gadalla 2010 ; Trottier et al., in press ; van der Bruggen and Grubb 2014 ). Prior comparisons of students and MTurkers in studies unrelated to rape perceptions show differences in population representativeness and data quality (Berinsky et al. 2012 ; Casler et al. 2013 ; Fleischer et al. 2015 ; Kees et al. 2017 ; Levay et al. 2016 ). Yet no study has comprehensively compared their strengths, weaknesses, or variation in findings in vignette experiments on rape perceptions. Furthermore, Bornstein et al. ( 2017 ) compared the effect of sample type—student or community—on jury verdicts, but they included both criminal and civil cases. Weigold and Weigold ( 2022 ) compared demographics, attention, time to completion, and social attitudes (e.g., egalitarianism) across students, MTurker nonstudents, and MTurker students, but they did not examine RMA or rape perceptions, the overall amount of usable data produced, or the cost and convenience of collecting the samples. Our study addressed these gaps and offered methodological recommendations regarding cost, convenience, data quality, and generalizability. We found MTurk more expensive but far more efficient than student sampling. Both produced comparable data quality, though students were more likely to pass manipulation checks and return usable surveys. Samples also differed demographically, as well as in RMA and victim blame. Based on these findings, we provide recommendations for researchers conducting experimental rape perception studies, particularly those focused on mock jury perceptions. 5.1. Implications and Recommendations First, while the extant literature on RMA and rape perceptions is vast, the methodological variations, strengths, weaknesses, and limitations of this literature have not yet been explored systematically. Periodic meta‐analyses and systematic reviews have focused on confirming specific relationships and outcomes of interest, such as documenting the demographic characteristics (e.g., gender) and attitudes (e.g., sexism) associated with RMA and victim blaming (Gravelin et al. 2019 ; Persson and Dhingra 2022 ; Suarez and Gadalla 2010 ; Trottier et al., in press ; van der Bruggen and Grubb 2014 ), or the effects of RMA on rape perceptions (Gravelin et al. 2019 ; Grubb and Turner 2012 ; Hockett et al. 2016 ; Persson and Dhingra 2022 ; Trottier et al., in press ; van der Bruggen and Grubb 2014 ), police decisions (Sleath and Bull 2017 ), and jury verdicts (Dinos et al. 2015 ; Leverick 2020 ). However, these studies have failed to describe how studies' methodologies influence these relationships, such as if sample size, type, or recruitment strategy influence findings. Nor have meta‐analyses and systematic reviews documented differences in study cost, data quality, or generalizability across study designs. Experienced researchers may understand, anecdotally, the pros and cons of different sampling strategies, including the cost, data quality, and generalizability of data collected from different convenience samples. But to our knowledge, our study is the first to document these concerns and to offer guidance to rape perception researchers. While our study offers some evidence that convenience sample type matters, we believe a more systematic review of study methodologies is warranted. We recommend that researchers carefully and consistently document the cost and data quality of their study designs, and we hope that future researchers conduct a systematic review and meta‐analysis detailing these characteristics. Second, researchers must recognize that sample source affects findings. Students exhibited lower RMA and victim blame than MTurkers. If we were to draw conclusions based only on students, we would argue that rape supportive attitudes are declining, or that they will not strongly influence jurors in sexual assault trials. We might further argue that victim characteristics, like gender, sexual orientation, and race, will not influence jurors' perceptions of the victim or case, which could encourage police and prosecutors to pursue cases involving victims from minoritized groups, including men and transgender people. Examining only MTurkers, however, we may draw different conclusions: RMA persists and is likely to impact jury perceptions of sexual assault; victim race influences jurors' assessments of victim blame and may reduce the likelihood of conviction in cases involving some victims. These differences show that sample source influences which factors appear consequential for prosecutors and jurors. Third, researchers should consider study goals. Those interested in describing attitudes among the general public may prefer MTurk, as MTurkers better matched national demographics (Hanson 2024 ; Pew Research Center 2016 ; U.S. Census Bureau 2021 ). By contrast, researchers interested in specific jury pools may prefer students at geographically relevant universities. For example, our Southwestern student sample more closely resembled the regional racial composition than the MTurk sample. Scholars examining race in local jury contexts, or focusing on young adults' RMA and rape perceptions, may find students more appropriate. The study's goals and population of interest, therefore, should guide researchers' sampling choices. Fourth, because conclusions depend on sample source, generalizability and replication are imperative. A comprehensive picture of rape perceptions—and jury perceptions more broadly—will only emerge through consistent replication. A replication crisis has been identified across disciplines, with implications for courtroom actors and decisions (Bardsley 2018 ; Chin 2014 ). Replicating rape perception studies with diverse samples, and evaluating them through systematic reviews and meta‐analyses, is the best way to identify causal relationships that are robust across study designs. Additionally, we encourage greater attention to how methodological variations, particularly in sampling, shape research questions and findings. Individual studies rarely provide a complete picture, especially when based solely on student samples. Replications, alongside rigorous periodic reviews of both published and unpublished work that describe studies' methodologies are essential. Fifth, our findings indicate that both students and MTurkers produce data with significant quality concerns. Students were less likely to fail manipulation checks, resulting in a larger proportion of usable surveys. Yet students are not necessarily preferable. MTurk facilitates collection of larger samples more quickly, and although only 48% of responses in the current study were usable, the MTurk sample still produced nearly four times as many usable cases as the student sample. Furthermore, researchers can improve the quality of MTurk responses at the outset by using intermediate platforms like CloudResearch.com , which screens MTurkers to identify high‐quality participants and block workers hiding their location or with low HIT rates. However, tools like CloudResearch.com add additional costs to research. MTurk already may be cost prohibitive, especially for student researchers, early career scholars, or nonprofits without external funding. Even researchers with funding may decide that dropping over 50% of paid responses is not a very good return on their investment. Researchers deciding which sample to use, therefore, must balance these priorities. Finally, given ethical concerns with MTurk, as well as the high proportion of missing data produced by both samples, researchers may consider alternatives such as Prolific Academic. MTurk's viability may diminish amid evidence of low‐quality responses (Fleischer et al. 2015 ), bots (Griffin et al. 2022 ; Xu et al. 2022 ), and low pay standards (Millum and Garnett 2019 ; Williamson 2016 ). Both MTurkers and researchers may migrate to other platforms that, while more expensive, may generate better quality data and more representative samples. The methodological considerations outlined here can help researchers select the most appropriate source—student, MTurk, or other crowdsourcing platforms—based on budget, timeline, and goals. 5.2. Limitations, Future Directions, and Conclusions The current study compared two common convenience samples: students and MTurkers. Future work should assess other methods, including alternative crowdsourcing platforms (e.g., Prolific Academic), social media recruitment (e.g., X, TikTok), and traditional approaches (e.g., phone or mail surveys). Each method has distinct strengths and weaknesses across cost, convenience, data quality, and generalizability. Additionally, our study manipulated victim characteristics—gender, sexual orientation, and race—but other case factors such as intoxication, resistance, or victim–perpetrator relationship may also vary by sample source. Future research should examine the effects of these factors on blame attributions among students, MTurkers, and other samples. Neither sample in this study is fully generalizable to their broader populations. Our student sample's racial composition reflected the university's Southwestern location, and students elsewhere may differ. Likewise, our MTurk sample differed somewhat from a representative sample of MTurkers collected by Moss and colleagues in 2019 (published in 2023), which had fewer men, more Latinx/Hispanic people, and more people without college degrees than our MTurk sample. These differences may reflect self‐selection bias, perhaps exacerbated by low compensation. Standardizing hourly pay across studies could reduce this bias; future work could test this by comparing MTurk samples on unrelated topics under high versus low pay conditions. Finally, we did not assess if respondents used artificial intelligence (AI) or robotic technologies to complete the survey. Relatedly, we were unable to detect or prevent the same MTurkers from completing the survey more than once using different worker IDs; a person could operate many different accounts, with different IDs, which would enable them or bots to complete the same survey multiple times. Specifically, researchers relying on social media and personal networks to administer online surveys have reported a high proportion of bot‐generated responses (Griffin et al. 2022 ; Xu et al. 2022 ). While the age of the data used in this study predates AI tools like ChatGPT, and the use of MTurk rather than social media to recruit respondents may have reduced the prevalence of bots, it is possible that at least some responses were generated by bots. Furthermore, as AI and bot use expands, online surveys may become increasingly vulnerable to fabricated data. Future research on sample sources should incorporate checks for AI‐ or bot‐generated responses. Using tools such as CloudResearch.com , which interfaces directly with MTurk, can prevent fraudulent responses and improve the overall quality of data. While the current study did not use this tool—we did not know about it at the time, and its cost would have exceeded the primary author's resources—researchers should consider all tools available to them in order to maximize data quality within the constraints of their resources. In conclusion, our study offers useful guidance to researchers studying rape perceptions, particularly regarding sample selection. Regarding data quality, students performed slightly better than MTurkers on manipulation checks—and due to this they produced proportionately more usable data—but the samples were otherwise comparable on data quality metrics. However, the MTurk sample was more convenient and efficient to collect, albeit much more costly, than the student sample. Furthermore, the demographic composition of both samples varied in important ways from the adult U.S. population, limiting generalizability of findings. However, the student sample seemed to resemble the specific racial and ethnic composition of the geographic region in which the university is located. Researchers trying to generalize their findings to a specific group—such as a geographically specific jury pool or young adults—may prefer a student sample drawn from a university within or near the relevant jurisdiction. Researchers, therefore, should select the sample that best accommodates their timeline, funding, study goals, and population of interest. Most importantly, MTurkers had higher RMA and blamed victims more than students. Additionally, the experiment yielded different findings in the student and MTurk samples. These findings suggest that rape perception studies relying on students may underestimate the continued pervasiveness of RMA and the problems it poses in jury trials, which underlines the importance of replicating studies using varied sample sources before asserting real world implications. Ethics Statement The original study protocol was reviewed by the Institutional Review Board at Arizona State University and continued use of the data was approved by the Institutional Review Board at the University of Arkansas at Little Rock. Conflicts of Interest The authors declare no conflicts of interest. Supporting information Supporting Information S1 BSL-44-249-s001.docx (30.1KB, docx) Acknowledgments This original data collection was funded in part by an undergraduate honors thesis grant awarded to Taylor Payne by the Barrett Honors College at Arizona State University in 2021. We would like to acknowledge and thank Ms. Payne for generously contributing their grant to pay research participants. Preliminary data and analyses related to this paper were presented at the Western Society of Criminology's annual conference in February 2023. The presentation described data collection methods and preliminary findings. Endnotes 1 State courts compile jury pools from voter registration and driver's license records, which often exclude nonvoters and nondrivers—systematically underrepresenting people of color and low‐income individuals (Equal Justice Initiative 2021 ). 2 Experimental vignette surveys commonly include manipulation checks, which are questions assessing if the respondent correctly noticed a manipulated element, such as the race of the victim. Long surveys and those using many Likert scale items may also include attention checks, or questions telling the respondent to select a particular response option. These attention check questions ensure that respondents actually read the questions rather than responding randomly or mindlessly. 3 In total, each participant read two vignettes, a date rape scenario that varied rape myth factors (e.g., “precipitation” or motives to lie) and an acquaintance rape that varied credibility issues (e.g., criminal record, substance abuse). Each vignette was followed by questions assessing perceptions of victim and perpetrator blame. The order of the vignettes was randomized to reduce priming effects. Blame attributions related to each scenario were always intended to be assessed separately, as if running two separate studies concurrently. Hence, the outcomes explored in this study focus on the perceptions of the date rape scenario only. 4 The survey used a Latin Square design to control for order effects of the vignettes and RMA scale (Bradley 1958 ). The order in which respondents read the vignettes and the RMA scale was randomized. Additionally, the order of the vignettes was randomized to control for stimulus order effects. 5 The original variable was highly skewed (mean = 53 min, SD = 549 min, median = 10.4 min, and a range of 0.07–15726.4 min) likely due to some respondents leaving their computer on or survey open while stepping away, which would artificially increase the recorded time in the survey. We replaced the top 5% of responses with the 95th percentile score of 34.75. The resulting variable was normally distributed (mean = 12.3, SD = 8.6, median = 10.4, range = 0.07–34.75). 6 We deemed surveys completed in less than 5 minutes unusable due to survey length and complexity. Surveys completed in less than 5 min were much more likely to fail attention checks compared to those completed in 5 min or more ( X 2 = 762.88, p = 0.000). 7 The primary author, Dr. Suzanne St. George, targeted instructors teaching multiple courses and courses with large class sizes. They also contacted instructors in varied disciplines, rather than only their own discipline (criminal justice). Courses targeted included: introduction to criminal justice, introduction to psychology, introduction to biology, introduction to physics, and tourism and management, among others. 8 Some instructors may have offered extra credit to students for participating in the study. Instructors were neither encouraged nor discouraged from offering extra credit, and whether they did so was not tracked. 9 Dr. Suzanne St. George paid $2035 of their own money to fund the study, while the remaining $1000 came from a grant from Arizona State University to support undergraduate research. The decision to spread data collection over 6 months was meant to reduce the financial burden of data collection for Dr. St. George. While $2000 may not seem like a lot of money to some, it may be cost prohibitive, particularly for students conducting research. Spreading the data collection out over time alleviated some of this burden on Dr. St. George, who at the time was a graduate research assistant. 10 All MTurkers were paid, including those that failed attention checks or skipped questions. While withholding payment may be a means to ensure higher quality responses, this seemed coercive and is inconsistent with IRB approved verbiage in the assent form stating there will be no penalty to respondents who skip questions or leave the survey unfinished. 11 In the IRB protocol, Dr. Suzanne St. George justified low MTurker pay as a safeguard against coercion, noting that high compensation might pressure participation on a sensitive topic. Because MTurkers can choose from many studies, those uninterested could easily select other tasks, further reducing coercion risk. Ethical debates, however, have shifted toward concerns about exploitation, with some arguing that reliance on MTurk as income warrants higher pay (Millum and Garnett 2019 ; Williamson 2016 ). Others note MTurkers' income distribution resembles the U.S. population, with most using it part‐time or for paid leisure (Moss et al. 2023 ), suggesting few depend on it as a primary income source. 12 Supporting Information S1 : Table S3 displays regression models of victim blame, while controlling for respondent characteristics and the experimental variables. In the regression models, victim gender remained statistically significantly associated with victim blame among students and MTurkers, but neither sexual orientation nor race were associated in either group. Data Availability Statement Data may be available from Dr. Suzanne St. George, upon request. References Aborisade, R. A. , Adegoke N., Adeleke O. A., et al. 2024. “Policing Rape and Serious Sexual Offenses in Nigeria: Officers’ Experiences and Appraisal of Police Investigative Approaches.” Police Practice and Research 25, no. 3: 251–268. 10.1080/15614263.2023.222870. [ DOI ] [ Google Scholar ] Alderden, M. A. , and Ullman S. E.. 2012. “Creating a More Complete and Current Picture: Examining Police and Prosecutor Decision‐Making When Processing Sexual Assault Cases.” Violence Against Women 18, no. 5: 525–551. 10.1177/1077801212453867. [ DOI ] [ PubMed ] [ Google Scholar ] Amazon Mechanical Turk . n.d. Amazon Mechanical Turk: Access a Global, on‐Demand, 24 × 7 Workforce. mturk.com . [ Google Scholar ] Aosved, A. C. , and Long P. J.. 2006. “Co‐Occurrence of Rape Myth Acceptance, Sexism, Racism, Homophobia, Ageism, Classism, and Religious Intolerance.” Sex Roles 55, no. 7: 481–492. 10.1007/s11199-006-9101-4. [ DOI ] [ Google Scholar ] Aruguete, M. S. , Huynh H., Browne B. L., Jurs B., Flint E., and McCutcheon L. E.. 2019. “How Serious Is the ‘Carelessness’ Problem on Mechanical Turk?” International Journal of Social Research Methodology 22, no. 5: 441–449. 10.1080/13645579.2018.1563966. [ DOI ] [ Google Scholar ] Bardsley, N.

2018. “What Lessons Does the ‘Replication Crisis’ in Psychology Hold for Experimental Economics?” In Handbook of Psychology and Economic Behaviour. 2nd ed. Cambridge Handbooks in Psychology. Cambridge University Press: ISBN 9781107161399. http://centaur.reading.ac.uk/69874/ . [ Google Scholar ] Behrend, T. S. , Sharek D. J., Meade A. W., and Wiebe E. N.. 2011. “The Viability of Crowdsourcing for Survey Research.” Behavior Research Methods 43, no. 3: 800–813. 10.3758/s13428-011-0081-0. [ DOI ] [ PubMed ] [ Google Scholar ] Berinsky, A. J. , Huber G. A., and Lenz G. S.. 2012. “Evaluating Online Labor Markets for Experimental Research: Amazon.com’s Mechanical Turk.” Political Analysis 20, no. 3: 351–368. 10.1093/pan/mpr057. [ DOI ] [ Google Scholar ] Bornstein, B. H. , Golding J. M., Neuschatz J., et al. 2017. “Mock Juror Sampling Issues in Jury Simulation Research: A Meta‐Analysis.” Law and Human Behavior 41, no. 1: 13–28. https://heinonline.org/HOL/P?h=hein.journals/lwhmbv41&i=13 . [ DOI ] [ PubMed ] [ Google Scholar ] Bornstein, B. H. , and Kleynhans A. J.. 2018. “The Evolution of Jury Research Methods: From Hugo Munsterberg to the Modern Age.” Denver Law Review 96: 813–840. https://digitalcommons.du.edu/dlr/vol96/iss4/8/ . [ Google Scholar ] Bradley, J. V.

1958. “Complete Counterbalancing of Immediate Sequential Effects in a Latin Square Design.” Journal of the American Statistical Association 53, no. 282: 525–528. 10.1080/01621459.1958.10501456. [ DOI ] [ Google Scholar ] Brown, J. D.

2008. “Effect Size and Eta Squared.” Shiken: JALT Testing and Evaluation Newsletter 12, no. 2: 38–43. https://teval.jalt.org/test/bro_28.htm . [ Google Scholar ] Buchheit, S. , Dalton D. W., Pollard T. J., and Stinson S. R.. 2019. “Crowdsourcing Intelligent Research Participants: A Student Versus MTurk Comparison.” Behavioral Research in Accounting 31, no. 2: 93–106. 10.2308/bria-52340. [ DOI ] [ Google Scholar ] Burt, M. R.

1980. “Cultural Myths and Supports for Rape.” Journal of Personality and Social Psychology 38, no. 2: 217–230. 10.1037/0022-3514.38.2.217. [ DOI ] [ PubMed ] [ Google Scholar ] Campbell, R. , and Fehler‐Cabral G.. 2018. “Why Police ‘Couldn’t or Wouldn’t’ Submit Sexual Assault Kits for Forensic DNA Testing: A Focal Concerns Theory Analysis of Untested Rape Kits.” Law & Society Review 52, no. 1: 73–105. 10.1111/lasr.12310. [ DOI ] [ Google Scholar ] Casler, K. , Bickel L., and Hackett E.. 2013. “Separate but Equal? A Comparison of Participants and Data Gathered via Amazon’s MTurk, Social Media, and Face‐to‐Face Behavioral Testing.” Computers in Human Behavior 29, no. 6: 2156–2160. 10.1016/j.chb.2013.05.009. [ DOI ] [ Google Scholar ] Chandler, J. , Mueller P., and Paolacci G.. 2014. “Non Naïveté Among Amazon Mechanical Turk Workers: Consequences and Solutions for Behavioral Researchers.” Behavior Research Methods 46, no. 1: 112–130. 10.3758/s13428-013-0365-7. [ DOI ] [ PubMed ] [ Google Scholar ] Chandler, J. , and Shapiro D.. 2016. “Conducting Clinical Research Using Crowdsourced Convenience Samples.” Annual Review of Clinical Psychology 12, no. 1: 53–81. 10.1146/annurev-clinpsy-021815-093623. [ DOI ] [ PubMed ] [ Google Scholar ] Chin, J. M.

2014. “Psychological Science’s Replicability Crisis and What It Means for Science in the Courtroom.” Psychology, Public Policy, and Law 20, no. 3: 225–238. 10.1037/law0000012. [ DOI ] [ Google Scholar ] Chmielewski, M. , and Kucker S. C.. 2020. “An MTurk Crisis? Shifts in Data Quality and the Impact on Study Results.” Social Psychological and Personality Science 11, no. 4: 464–473. 10.1177/194855061987514. [ DOI ] [ Google Scholar ] Coble, S.

2022. “Intersections of Racism and Sexism in Rape Myth Research: Exploring How Race Conditions the Effects of Rape Myths on Rape Perceptions and Criminal Justice Responses.” Doctoral dissertation, Arizona State University, no. 29167667. Davies, M. , Gilston J., and Rogers P.. 2012. “Examining the Relationship Between Male Rape Myth Acceptance, Female Rape Myth Acceptance, Victim Blame, Homophobia, Gender Roles, and Ambivalent Sexism.” Journal of Interpersonal Violence 27, no. 14: 2807–2823. 10.1177/0886260512438281. [ DOI ] [ PubMed ] [ Google Scholar ] Devine, D. J.

2012. Jury Decision Making: The State of the Science. NYU Press. [ Google Scholar ] Dewald, S. , and Lorenz K.. 2022. “Lying About Sexual Assault: A Qualitative Study of Detective Perspectives on False Reporting.” Policing and Society 32, no. 2: 179–199. 10.1080/10439463.2021.1893725. [ DOI ] [ Google Scholar ] Dickert, N. , and Grady C.. 1999. “What’s the Price of a Research Subject? Approaches to Payment for Research Participation.” New England Journal of Medicine 341, no. 3: 198–203. 10.1056/NEJM199907153410312. [ DOI ] [ PubMed ] [ Google Scholar ] Dinos, S. , Burrowes N., Hammond K., and Cunliffe C.. 2015. “A Systematic Review of Juries’ Assessment of Rape Victims: Do Rape Myths Impact on Juror Decision‐Making?” International Journal of Law, Crime and Justice 43, no. 1: 36–49. 10.1016/j.ijlcj.2014.07.001. [ DOI ] [ Google Scholar ] Dong, Y. , and Peng C. Y. J.. 2013. “Principled Missing Data Methods for Researchers.” SpringerPlus 2, no. 1: 222. 10.1186/2193-1801-2-222. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Equal Justice Initiative . 2021. Race and the Jury: Illegal Racial Discrimination in Jury Selection. https://eji.org/report/race‐and‐the‐jury/ . Fávero, M. , Del Campo A., Faria S., Moreira D., Ribeiro F., and Sousa‐Gomes V.. 2022. “Rape Myth Acceptance of Police Officers in Portugal.” Journal of Interpersonal Violence 37, no. 1–2: 659–680. 10.1177/0886260520916282. [ DOI ] [ PubMed ] [ Google Scholar ] Fleischer, A. , Mead A. D., and Huang J.. 2015. “Inattentive Responding in MTurk and Other Online Samples.” Industrial and Organizational Psychology 8, no. 2: 196–202. 10.1017/iop.2015.25. [ DOI ] [ Google Scholar ] Garza, A. D. , and Franklin C. A.. 2021. “The Effect of Rape Myth Endorsement on Police Response to Sexual Assault Survivors.” Violence Against Women 27, no. 3–4: 552–573. 10.1177/1077801220911460. [ DOI ] [ PubMed ] [ Google Scholar ] Gavin, S. , Kosaka R., Bangura M. D., Kruis N. E., and Rowland N. J.. in press. “Rape Myth Acceptance in the Criminal Justice System: Do Criminal Justice Decision‐Makers Endorse Rape Myths?” Journal of Interpersonal Violence: 08862605251363618. Advanced Online Publication. 10.1177/08862605251363618. [ DOI ] [ PubMed ] [ Google Scholar ] Gekoski, A. , Massey K., Allen K., et al. 2023. “A Lot of the Time It’s Dealing With Victims Who Don’t Want to Know, It’s all Made Up, or They’ve Got Mental Health’: Rape Myths in a Large English Police Force.” International Review of Victimology 30, no. 1: 3–24. 10.1177/02697580221142891. [ DOI ] [ Google Scholar ] Goodman, J. K. , Cryder C. E., and Cheema A.. 2013. “Data Collection in a Flat World: The Strengths and Weaknesses of Mechanical Turk Samples.” Journal of Behavioral Decision Making 26, no. 3: 213–224. 10.1002/bdm.1753. [ DOI ] [ Google Scholar ] Goodman, J. K. , and Paolacci G.. 2017. “Crowdsourcing Consumer Research.” Journal of Consumer Research 44, no. 1: 196–210. 10.1093/jcr/ucx047. [ DOI ] [ Google Scholar ] Graham, J. W.

2009. “Missing Data Analysis: Making It Work in the Real World.” Annual Review of Psychology 60, no. 1: 549–576. 10.1146/annurev.psych.58.110405.085530. [ DOI ] [ PubMed ] [ Google Scholar ] Gravelin, C. R. , Biernat M., and Bucher C. E.. 2019. “Blaming the Victim of Acquaintance Rape: Individual, Situational, and Sociocultural Factors.” Frontiers in Psychology 9: 2422. 10.3389/fpsyg.2018.02422. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Griffin, M. , Martino R. J., LoSchiavo C., et al. 2022. “Ensuring Survey Research Data Integrity in the Era of Internet Bots.” Quality and Quantity 55, no. 3: 1–12. 10.1007/s11135-021-01252-1. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Grubb, A. , and Turner E.. 2012. “Attribution of Blame in Rape Cases: A Review of the Impact of Rape Myth Acceptance, Gender Role Conformity and Substance Use on Victim Blaming.” Aggression and Violent Behavior 17, no. 5: 443–452. 10.1016/j.avb.2012.06.002. [ DOI ] [ Google Scholar ] Hanson, M.

2024. “College Enrollment & Student Demographic Statistics.” EducationData.org, December 21. https://educationdata.org/college‐enrollment‐statistics . [ Google Scholar ] Hauser, D. J. , and Schwarz N.. 2016. “Attentive Turkers: Mturk Participants Perform Better on Online Attention Checks Than Do Subject Pool Participants.” Behavior Research Methods 48, no. 1: 400–407. 10.3758/s13428-015-0578-z. [ DOI ] [ PubMed ] [ Google Scholar ] Helm, R. K.

2024. How Juries Work: And How They Could Work Better: [Google Books]. Hermolle, M. , Kent A., Locke A. J., and Andrews S. J.. 2024. “‘Are We Sure That He Knew That You Don’t Want to Have Sex?’: Discursive Constructions of the Suspect in Police Interviews With Rape Complainants.” Behavioral Sciences 14, no. 9: 837. 10.3390/bs14090837. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Hockett, J. M. , Smith S. J., Klausing C. D., and Saucier D. A.. 2016. “Rape Myth Consistency and Gender Differences in Perceiving Rape Victims: A Meta‐Analysis.” Violence Against Women 22, no. 2: 139–167. 10.1177/1077801215607359. [ DOI ] [ PubMed ] [ Google Scholar ] Kees, J. , Berry C., Burton S., and Sheehan K.. 2017. “An Analysis of Data Quality: Professional Panels, Student Subject Pools, and Amazon’s Mechanical Turk.” Journal of Advertising 46, no. 1: 141–155. 10.1080/00913367.2016.1269304. [ DOI ] [ Google Scholar ] Keller, S. R. , and Wiener R. L.. 2011. “What Are We Studying? Student Jurors, Community Jurors, and Construct Validity.” Behavioral Sciences & the Law 29, no. 3: 376–394. 10.1002/bsl.971. [ DOI ] [ PubMed ] [ Google Scholar ] Lattuca, L. R. , and Stark J. S.. 2009. Shaping the College Curriculum: Academic Plans in Context. 2nd ed. Jossey‐Bass. [ Google Scholar ] Lee, J. , Lee C., and Lee W.. 2012. “Attitudes Toward Women, Rape Myths, and Rape Perceptions Among Male Police Officers in South Korea.” Psychology of Women Quarterly 36, no. 3: 365–376. 10.1177/0361684311427538. [ DOI ] [ Google Scholar ] Lehmann, J. Y. K. , and Smith J. B.. 2013. A Multidimensional Examination of Jury Composition, Trial Outcomes, and Attorney Preferences. University of Houston: [Unpublished manuscript]. http://www.uh.edu/~jlehman2/papers/lehmann_smith_jurycomposition.pdf . [ Google Scholar ] Levay, K. E. , Freese J., and Druckman J. N.. 2016. “The Demographic and Political Composition of Mechanical Turk Samples.” Sage Open 6, no. 1: 1–17. 10.1177/2158244016636433. [ DOI ] [ Google Scholar ] Leverick, F.

2020. “What Do We Know About Rape Myths and Juror Decision Making?” International Journal of Evidence and Proof 24, no. 3: 255–279. 10.1177/1365712720923157. [ DOI ] [ Google Scholar ] Lonsway, K. A. , and Fitzgerald L. F.. 1994. “Rape Myths. In Review.” Psychology of Women Quarterly 18, no. 2: 133–164. 10.1111/j.1471-6402.1994.tb00448.x. [ DOI ] [ Google Scholar ] McMahon, S. , and Farmer G. L.. 2011. “An Updated Measure for Assessing Subtle Rape Myths.” Social Work Research 35, no. 2: 71–81. 10.1093/swr/35.2.71. [ DOI ] [ Google Scholar ] Millum, J. , and Garnett M.. 2019. “How Payment for Research Participation Can Be Coercive.” American Journal of Bioethics 19, no. 9: 21–31. 10.1080/15265161.2019.1630497. [ DOI ] [ PubMed ] [ Google Scholar ] Morabito, M. S. , Pattavina A., and Williams L. M.. 2019. “It all Just Piles Up: Challenges to Victim Credibility Accumulate to Influence Sexual Assault Case Processing.” Journal of Interpersonal Violence 34, no. 15: 3151–3170. 10.1177/0886260516669164. [ DOI ] [ PubMed ] [ Google Scholar ] Moss, A. J. , Rosenzweig C., Robinson J., Jaffe S. N., and Litman L.. 2023. “Is It Ethical to Use Mechanical Turk for Behavioral Research? Relevant Data From a Representative Survey of MTurk Participants and Wages.” Behavior Research Methods 55, no. 8: 4048–4067. 10.3758/s13428-022-02005-0. [ DOI ] [ PubMed ] [ Google Scholar ] Muehlenhard, C. L. , Peterson Z. D., Humphreys T. P., and Jozkowski K. N.. 2017. “Evaluating the One‐in‐Five Statistic: Women’s Risk of Sexual Assault While in College.” Journal of Sex Research 54, no. 4–5: 549–576. 10.1080/00224499.2017.1295014. [ DOI ] [ PubMed ] [ Google Scholar ] Mumford, E. A. , Potter S., Taylor B. G., and Stapleton J.. 2020. “Sexual Harassment and Sexual Assault in Early Adulthood: National Estimates for College and Non‐College Students.” Public Health Reports 135, no. 5: 555–559. 10.1177/0033354920946014. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] OpenAI . 2025. ChatGPT (April 2025 Version). https://openai.com/chatgpt . Paolacci, G. , and Chandler J.. 2014. “Inside the Turk: Understanding Mechanical Turk as a Participant Pool.” Current Directions in Psychological Science 23, no. 3: 184–188. 10.1177/0963721414531598. [ DOI ] [ Google Scholar ] Peer, E. , Vosgerau J., and Acquisti A.. 2014. “Reputation as a Sufficient Condition for Data Quality on Amazon Mechanical Turk.” Behavior Research Methods 46, no. 4: 1023–1031. 10.3758/s13428-013-0434-y. [ DOI ] [ PubMed ] [ Google Scholar ] Pennington, N. , and Hastie R.. 1991. “A Cognitive Theory of Juror Decision Making: The Story Model.” Cardozo Law Review 13: 519–558. https://heinonline.org/HOL/LandingPage?handle=hein.journals/cdozo13&div=30&id=&page= . [ Google Scholar ] Persson, S. , and Dhingra K.. 2022. “Attributions of Blame in Stranger and Acquaintance Rape: A Multilevel Meta‐Analysis and Systematic Review.” Trauma, Violence, & Abuse 23, no. 3: 795–809. 10.1177/1524838020977146. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] PettyJohn, M. E. , Cary K. M., and McCauley H. L.. 2023. “Rape Myth Acceptance in a Community Sample of Adult Women in the Post# MeToo Era.” Journal of Interpersonal Violence 38, no. 13–14: 8211–8234. 10.1177/08862605231153893. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Pew Research Center . 2016. “A Wider Ideological Gap Between More and less Educated Adults: Political Polarization Update.” April 26. https://www.pewresearch.org/politics/2016/04/26/a‐wider‐ideological‐gap‐between‐more‐and‐less‐educated‐adults/ . Ross, L.

2024. “Mock Juries, Real Trials: How to Solve (Some) Problems With Jury Science.” Journal of Law and Society 51, no. 3: 324–342. 10.1111/jols.12494. [ DOI ] [ Google Scholar ] Salerno‐Ferraro, A. C. , and Jung S.. 2022. “To Charge or Not to Charge? Police Decisions in Canadian Sexual Assault Cases and the Relevance of Rape Myths.” Police Practice and Research 23, no. 5: 539–552. 10.1080/15614263.2021.2002150. [ DOI ] [ Google Scholar ] Sharma, P. , and Hamilton G.. 2025. “Police Responses to Rape in Metropolitan India.” International Journal for Crime, Justice and Social Democracy 14, no. 3: 91–103. 10.5204/ijcjsd.3409. [ DOI ] [ Google Scholar ] Sinclair, O.

2022. “The Attrition Problem: The Role of Police Officer’s Decision Making in Rape Cases.” Journal of Investigative Psychology and Offender Profiling 19, no. 3: 135–150. 10.1002/jip.1601. [ DOI ] [ Google Scholar ] Sleath, E. , and Bull R.. 2015. “A Brief Report on Rape Myth Acceptance: Differences Between Police Officers, Law Students, and Psychology Students in the United Kingdom.” Violence & Victims 30, no. 1: 136–147. 10.1891/0886-6708.VV-D-13-00035. [ DOI ] [ PubMed ] [ Google Scholar ] Sleath, E. , and Bull R.. 2017. “Police Perceptions of Rape Victims and the Impact on Case Decision Making: A Systematic Review.” Aggression and Violent Behavior 34: 102–112. 10.1016/j.avb.2017.02.003. [ DOI ] [ Google Scholar ] Smith, S. M. , Roster C. A., Golden L. L., and Albaum G. S.. 2016. “A Multi‐Group Analysis of Online Survey Respondent Data Quality: Comparing a Regular USA Consumer Panel to MTurk Samples.” Journal of Business Research 69, no. 8: 3139–3148. 10.1016/j.jbusres.2015.12.002. [ DOI ] [ Google Scholar ] Spohn, C. , and Tellis K.. 2014. Policing and Prosecuting Sexual Assault: Inside the Criminal Justice System. Lynne Rienner Publishers. [ Google Scholar ] St. George, S.

2025. “Intersecting Rape Myths With Race: Examining Race‐and Ethnicity‐Specific Effects of Rape Myth Factors on Police Responses to Sexual Assault.” Justice Quarterly 42, no. 1: 90–119. 10.1080/07418825.2023.2263531. [ DOI ] [ Google Scholar ] St. George, S. , and Spohn C.. 2018. “Liberating Discretion: The Effect of Rape Myth Factors on Prosecutors’ Decisions to Charge Suspects in Penetrative and Non‐Penetrative Sex Offenses.” Justice Quarterly 35, no. 7: 1280–1308. 10.1080/07418825.2018.1529251. [ DOI ] [ Google Scholar ] St. George, S. , Verhagan M., and Spohn C.. 2022. “Detectives’ Descriptions of Their Responses to Sexual Assault Cases and Victims: Assessing the Overlap Between Rape Myths and Focal Concerns.” Police Quarterly 25, no. 1: 90–117. 10.1177/10986111211037592. [ DOI ] [ Google Scholar ] Suarez, E. , and Gadalla T. M.. 2010. “Stop Blaming the Victim: A Meta‐Analysis on Rape Myths.” Journal of Interpersonal Violence 25, no. 11: 2010–2035. 10.1177/0886260509354503. [ DOI ] [ PubMed ] [ Google Scholar ] Thelan, A. R. , and Meadows E. A.. 2022. “The Illinois Rape Myth Acceptance Scale—Subtle Version: Using an Adapted Measure to Understand the Declining Rates of Rape Myth Acceptance.” Journal of Interpersonal Violence 37, no. 19–20: NP17807–NP17833. 10.1177/08862605211030013. [ DOI ] [ PubMed ] [ Google Scholar ] Toich, M. J. , Schutt E., and Fisher D. M.. 2022. “Do You Get what You Pay For? Preventing Insufficient Effort Responding in MTurk and Student Samples.” Applied Psychology 71, no. 2: 640–661. 10.1111/apps.12344. [ DOI ] [ Google Scholar ] Trottier, D. , Laviolette V., Tuzi I., and Benbouriche M.. in press. “The Effect of Gender Role Expectations, Sexism, and Rape Myth Acceptance on the Social Perception of Sexual Violence: A Meta‐Analysis.” Trauma, Violence, & Abuse: 15248380251343190: Advanced online publication. 10.1177/15248380251343190. [ DOI ] [ PubMed ] [ Google Scholar ] United States Census Bureau . 2021. “American Community Survey: Narrative Profiles, 2017–2021.” https://www.census.gov/acs/www/data/data‐tables‐and‐tools/narrative‐profiles/2021/report.php?geotype=nation&usVal=us . van der Bruggen, M. , and Grubb A.. 2014. “A Review of the Literature Relating to Rape Victim Blaming: An Analysis of the Impact of Observer and Victim Characteristics on Attribution of Blame in Rape Cases.” Aggression and Violent Behavior 19, no. 5: 523–531. 10.1016/j.avb.2014.07.008. [ DOI ] [ Google Scholar ] Venema, R. M.

2019. “Making Judgments: How Blame Mediates the Influence of Rape Myth Acceptance in Police Response to Sexual Assault.” Journal of Interpersonal Violence 34, no. 13: 2697–2722. 10.1177/0886260516662437. [ DOI ] [ PubMed ] [ Google Scholar ] Weigold, A. , and Weigold I. K.. 2022. “Traditional and Modern Convenience Samples: An Investigation of College Student, Mechanical Turk, and Mechanical Turk College Student Samples.” Social Science Computer Review 40, no. 5: 1302–1322. 10.1177/08944393211006847. [ DOI ] [ Google Scholar ] Wiener, R. L. , Krauss D. A., and Lieberman J. D.. 2011. “Mock Jury Research: Where Do We Go From Here?” Behavioral Sciences & the Law 29, no. 3: 467–479. 10.1002/bsl.989. [ DOI ] [ PubMed ] [ Google Scholar ] Williamson, V.

2016. “On the Ethics of Crowdsourced Research.” PS: Political Science & Politics 49, no. 1: 77–81. 10.1017/S104909651500116X. [ DOI ] [ Google Scholar ] Xu, Y. , Pace S., Kim J., et al. 2022. “Threats to Online Surveys: Recognizing, Detecting, and Preventing Survey Bots.” Social Work Research 46, no. 4: 343–350. 10.1093/swr/svac023. [ DOI ] [ Google Scholar ] Associated Data This section collects any data citations, data availability statements, or supplementary materials included in this article. Supplementary Materials Supporting Information S1 BSL-44-249-s001.docx (30.1KB, docx) Data Availability Statement Data are not publicly available, but may be available from Dr. Suzanne St. George, upon request. Data may be available from Dr. Suzanne St. George, upon request. Articles from Behavioral Sciences & the Law are provided here courtesy of Wiley ACTIONS View on publisher site PDF (825.0 KB) Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top

Record · ID 25841 · SHA-256 d6b10b851ac9f1fc
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.