DOI QR코드

DOI QR Code

Group Differences in Memory Performance Among Amazon Mechanical Turk Masters, Regular Workers, and Offline Participants

Amazon Mechanical Turk 마스터, 일반 참가자, 오프라인 참가자 집단의 기억 수행 차이

  • Received : 2024.07.24
  • Accepted : 2024.09.26
  • Published : 2024.12.31

Abstract

The online crowdsourcing platform Amazon Mechanical Turk (MTurk) assigns a "master" qualification to workers who have outstanding task performance records. However, prior research comparing MTurk's master and regular workers has shown inconsistent results regarding actual performance differences between these groups. Furthermore, studies have used a survey method and research comparing cognitive task performance between MTurk masters and regular workers is still limited. The current study compared the performance of MTurk masters, regular workers, and offline-recruited university students using a visual recognition memory task. Results showed comparable memory performances between MTurk masters and offline participants. However, MTurk regular workers exhibited a different pattern of results from those of the masters and offline participants. Consistent results were found after excluding low-performing participants from each group. These findings suggest that appropriately screened online participants can effectively replicate results from traditional offline experiments. However, the results also underscore that online crowdsourcing platforms such as MTurk are made up of heterogeneous participant groups, suggesting that study outcomes may vary depending on participant selection criteria.

온라인 크라우드소싱 플랫폼인 Amazon Mechanical Turk(MTurk)은 뛰어난 과제 수행 기록을 가진 참가자들에게 마스터 등급을 부여한다. 그러나 MTurk의 마스터 참가자와 일반 참가자를 비교한 선행 연구들은 두 집단이 실제로 수행의 차이를 보이는가에 대해 일관되지 않은 결과를 보고했다. 또한 선행 연구들은 대부분 설문 조사 방식을 사용했으며 MTurk의 마스터와 일반 참가자의 인지 과제 수행 능력을 비교한 연구는 부족한 상황이다. 본 연구는 시각 기억 재인 과제를 사용하여 MTurk 마스터 및 일반 참가자와 오프라인에서 모집한 대학생 참가자 집단의 수행을 비교했다. 연구 결과, MTurk 마스터 참가자와 오프라인 참가자는 동일한 수준의 기억 수행을 보였다. 그러나 MTurk 일반 참가자의 기억 과제 수행은 마스터와 오프라인 참가자 집단의 결과와 차이를 보였다. 각 집단에서 기억 과제 정확률이 낮은 참가자를 제외한 후에도 동일한 결과가 나타났다. 이러한 결과는 온라인에서 참가자 집단을 적절히 선발하면 기존의 오프라인 실험 결과를 잘 재현할 수 있음을 보여준다. 동시에 본 연구의 결과는 온라인 크라우드소싱 플랫폼의 참가자 집단이 균일하지 않으며, 집단 선정 방식에 따라 연구의 결과가 다르게 나타날 수 있음을 시사한다.

Keywords

References

  1. Aguinis, H., Villamor, I., & Ramani, R. S. (2021). MTurk research: review and recommendations. Journal of Management, 47(4), 823-837. DOI: 10.1177/0149206320969787
  2. Albert, D. A., & Smilek, D. (2023). Comparing attentional disengagement between prolific and MTurk samples. Scientific Reports, 13(1), 20574. DOI: 10.1038/s41598-023-46048-5
  3. Amazon. (2011). Requester best practices guide. Amazon Web Services. Retrieved from http://mturkpublic.s3.amazonaws.com/docs/MTURK_BP.pdf
  4. Amazon. (2024). Amazon Mechanical Turk: Requester UI Guide. Amazon Web Services. Retrieved from https://docs.aws.amazon.com/pdfs/AWSMechTurk/latest/RequesterUI/amt-ui.pdf
  5. Arechar, A. A., & Rand, D. G. (2021). Turking in the time of COVID. Behavior Research Methods, 53(6), 2591-2595. DOI: 10.3758/s13428-021-01588-4
  6. Aruguete, M. S., Huynh, H., Browne, B. L., Jurs, B., Flint, E., & McCutcheon, L. E. (2019). How serious is the 'carelessness' problem on Mechanical Turk? international Journal of Social Research Methodology, 22(5), 441-449. DOI: 10.1080/13645579.2018.1563966
  7. Bakker, M., Van Dijk, A., & Wicherts, J. M. (2012). The rules of the game called psychological science. Perspectives on Psychological Science, 7(6), 543-554. DOI: 10.1177/1745691612459060
  8. Brascamp, J. W. (2021). Controlling the spatial dimensions of visual stimuli in online experiments. Journal of Vision, 21(8), 19-19. DOI: 10.1167/jov.21.8.19
  9. Buhrmester, M., Kwang, T., & Gosling, S. D. (2011). Amazon's Mechanical Turk: a new source of inexpensive, yet high-quality, data? Perspectives on Psychological Science, 6(1), 3-5. DOI: 10.1177/1745691610393980
  10. Buhrmester, M. D., Talaifar, S., & Gosling, S. D. (2018). An evaluation of Amazon's Mechanical Turk, its rapid rise, and its effective use. Perspectives on Psychological Science, 13(2), 149-154. DOI: 10.1177/174569161770651
  11. Chandler, J., Paolacci, G., Peer, E., Mueller, P., & Ratliff, K. A. (2015). Using nonnaive participants can reduce effect sizes. Psychological Science, 26(7), 1131-1139. DOI: 10.1177/095679761558511
  12. De Man, J., Campbell, L., Tabana, H., & Wouters, E. (2021). The pandemic of online research in times of COVID-19. BMJ Open, 11(2), e043866. DOI: 10.1136/bmjopen-2020-043866
  13. Douglas, B. D., Ewell, P. J., & Brauer, M. (2023). Data quality in online human-subjects research: comparisons between MTurk, Prolific, CloudResearch, Qualtrics, and SONA. PloS One, 18(3), e0279720. DOI: 10.1371/journal.pone.0279720
  14. Germine, L., Nakayama, K., Duchaine, B. C., Chabris, C. F., Chatterjee, G., & Wilmer, J. B. (2012). Is the Web as good as the lab? Comparable performance from Web and lab in cognitive/ perceptual experiments. Psychonomic Bulletin & Review, 19(5), 847-857. DOI: 10.3758/s13423-012-0296-9
  15. Germine, L. T., Duchaine, B., & Nakayama, K. (2011). Where cognitive development and aging meet: face learning ability peaks after age 30. Cognition, 118(2), 201-210. DOI: 10.1016/j.cognition.2010.11.002
  16. Goodman, J. K., Cryder, C. E., & Cheema, A. (2013). Data collection in a flat world: the strengths and weaknesses of Mechanical Turk samples. Journal of Behavioral Decision Making, 26(3), 213-224. DOI: 10.1002/bdm.1753
  17. Gosling, S. D., Vazire, S., Srivastava, S., & John, O. P. (2004). Should we trust web-based studies? a comparative analysis of six preconceptions about internet questionnaires. American Psychologist, 59(2), 93. DOI: 10.1037/0003-066X.59.2.93
  18. Hauser, D. J., & Schwarz, N. (2016). Attentive Turkers: MTurk participants perform better on online attention checks than do subject pool participants. Behavior Research Methods, 48(1), 400-407. DOI: 10.3758/s13428-015-0578-z
  19. Henrich, J., Heine, S. J., & Norenzayan, A. (2010). The weirdest people in the world?. Behavioral and Brain Sciences, 33(2-3), 61-83. DOI: 10.1017/S0140525X0999152X
  20. JASP Team. (2023). JASP (Version 0.17.2)[Computer software]. https://jasp-stats.org/
  21. Jeong, S. K. (2023a). Perceived image size modulates visual memory. Psychonomic Bulletin & Review, 30(6), 2282-2288. DOI: 10.3758/s13423-023-02313-2
  22. Jeong, S. K. (2023b). Cross-cultural consistency of image memorability. Scientific Reports, 13(1), 12737. DOI: 10.1038/s41598-023-39988-5
  23. Kim, B., Baek, H., Lee, Y., & Choi, W. (2021). Replication of word predictability effects using a web-based self-paced reading task. Korean Journal of Cognitive and Biological Psychology, 33(2), 87-93. DOI: 10.22172/cogbio.2021.33.2.001
  24. Lee, S., Nam, Y.-E., & Lee, Y. (2021). Studying attention with the web-based on-line experiment: is an online experiment as effective as a laboratory experiment? Journal of The Korean Data Analysis Society, 23(3), 1355-1368. DOI: 10.37727/jkdas.2021.23.3.1355
  25. Loepp, E., & Kelly, J. T. (2020). Distinction without a difference? An assessment of MTurk Worker types. Research & Politics, 7(1), 2053168019901185. DOI: 10.1177/205316801990118
  26. Lovett, M., Bajaba, S., Lovett, M., & Simmering, M. J. (2017). Data quality from crowdsourced surveys: a mixed method inquiry into perceptions of Amazon's Mechanical Turk Masters. Applied Psychology, 67(2), 339-366. DOI: 10.1111/apps.12124
  27. Masarwa, S., Kreichman, O., & Gilaie-Dotan, S. (2022). Larger images are better remembered during naturalistic encoding. Proceedings of the National Academy of Sciences of the United States of America, 119(4). e2119614119. DOI: 10.1073/pnas.2119614119
  28. Mason, W., & Suri, S. (2012). Conducting behavioral research on Amazon's Mechanical Turk. Behavior Research Methods, 44(1), 1-23. DOI: 10.3758/s13428-011-0124-6
  29. Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. Science, 349(6251), aac4716. DOI: 10.1126/science.aac4716
  30. Peer, E., Vosgerau, J., & Acquisti, A. (2014). Reputation as a sufficient condition for data quality on Amazon Mechanical Turk. Behavior Research Methods, 46(4), 1023-1031. DOI: 10.3758/s13428-013-0434-y
  31. Peirce, J., Gray, J. R., Simpson, S., MacAskill, M., Hochenberger, R., Sogo, H., Kastman, E., & Lindelov, J. K. (2019). PsychoPy2: Experiments in behavior made easy. Behavior Research Methods, 51(1), 195-203. DOI: 10.3758/s13428-018-01193-y
  32. Porfido, C. L., Cox, P. H., Adamo, S. H., & Mitroff, S. R. (2019). Recruiting from the shallow end of the pool: differences in cognitive and compliance measures for subject pool participants based on enrollment time across an academic term. Visual Cognition, 28(1), 1-9. DOI: 10.1080/13506285.2019.1702602
  33. Robinson, J., Rosenzweig, C., Moss, A. J., & Litman, L. (2019). Tapped out or barely tapped? Recommendations for how to harness the vast and largely unused potential of the Mechanical Turk participant pool. PloS One, 14(12), e0226394. DOI: 10.1371/journal.pone.0226394
  34. Rouse, S. V. (2020). Reliability of MTurk data from masters and workers. Journal of Individual Differences, 41(1), 30-36. DOI: 10.1027/1614-0001/a000300
  35. Stewart, N., Chandler, J., & Paolacci, G. (2017). Crowdsourcing samples in cognitive science. Trends in Cognitive Sciences, 21(10), 736-748. DOI: 10.1016/j.tics.2017.06.007
  36. Tae, J., Kim, T., & Choi, W. (2021). Reinvestigating the phonological and orthographic priming effects with the web-based experiment. Korean Journal of Linguistics, 46(4), 1223-1250. DOI: 10.18855/lisoko.2021.46.4.012
  37. Trenge, C., & Griffith, J. D. (2024). Master Turkers: An assessment of data quality. Measurement Instruments for the Social Sciences, 6, 1-18. DOI: 10.5964/miss.13619
  38. Yonelinas, A. P., Ramey, M. M., Riddell, C., Kahana, M., & Wagner, A. (2022). Recognition memory: The role of recollection and familiarity. In M. J. Kahana & A. D. Wagner (Eds.), The Oxford handbook of human memory. Oxford University Press.