Education, Science, Technology, Innovation and Life
Open Access
Sign In

Silicon Samples: A Review and Outlook on Large Language Models Simulating Human Respondents

Download as PDF

DOI: 10.23977/jaip.2026.090119 | Downloads: 4 | Views: 36

Author(s)

Yuyi Zhang 1

Affiliation(s)

1 School of Business Administration, Guizhou University of Finance and Economics, Guiyang, Guizhou, 550025, China

Corresponding Author

Yuyi Zhang

ABSTRACT

The use of large language models (LLMs) to simulate human respondents (silicon samples), as an emerging research topic, is attracting growing scholarly attention. By reviewing 25 representative studies from both domestic and international literature, this paper offers a definition of LLM-based simulation of human samples and examines its foundations, characteristics, and purposes. It summarizes the applications of this method across four major domains—psychometrics and machine psychology, consumer and market research, experimental simulation and digital twin construction, and the simulation of political opinion and social sentiment. It further reviews the existing evidence regarding the method's validity, along with the attendant debates, along three dimensions: psychometric validity, group fidelity and context dependence, and ethical and epistemic justice risks. Finally, the paper synthesizes a research framework for LLM-based simulation of human samples and discusses future research directions, with the aim of providing a reference for research and applied practice in this field.

KEYWORDS

Large language models; Silicon samples; Algorithmic Fidelity; Psychometrics and Machine Psychology; Consumer and Market Research; Digital Twins; Political Opinion Simulation

CITE THIS PAPER

Yuyi Zhang. Silicon Samples: A Review and Outlook on Large Language Models Simulating Human Respondents. Journal of Artificial Intelligence Practice (2026). Vol. 9, No. 1, 179-186. DOI: http://dx.doi.org/10.23977/jaip.2026.090119.

REFERENCES

[1] Bisbee, J., Clinton, J. D., Dorff, C., Kenkel, B., & Larson, J. M. (2024). Synthetic replacements for human survey data? The perils of large language models. Political Analysis, 32, 401–416. https://doi.org/10.1017/pan.2024.5
[2] Argyle, L. P., Busby, E. C., Fulda, N., Gubler, J. R., Rytting, C., & Wingate, D. (2023). Out of one, many: Using language models to simulate human samples. Political Analysis, 31(3), 337–351. https://doi.org/10.1017/pan.2023.2
[3] Toubia, O., Gui, G. Z., Peng, T., Merlau, D. J., Li, A., & Chen, H. (2025). Database report: Twin-2K-500: A data set for building digital twins of over 2,000 people based on their answers to over 500 questions. Marketing Science. https://doi.org/10.1287/mksc.2025.0262
[4] Wang, A., Morgenstern, J., & Dickerson, J. P. (2025). Large language models that replace human participants can harmfully misportray and flatten identity groups. Nature Machine Intelligence, 7(3), 400–411. https://doi.org/10.1038/s42256-025-00986-z
[5] Che, S., Zhu, M., Zhang, S., Jung, H. S., Lee, H., Wang, Z., & Miller, L. (2026). Simulating the people's voice: Leveraging algorithmic fidelity to assess ChatGPT's performance in modeling public opinion on Chinese government policies. Information Processing and Management, 63(1), 104567. https://doi.org/10.1016/j.ipm.2025.104567
[6] Lyman, A., Hepner, B., Argyle, L. P., Busby, E. C., Gubler, J. R., & Wingate, D. (2025). Balancing large language model alignment and algorithmic fidelity in social science research. Sociological Methods & Research, 54(3), 1110–1155. https://doi.org/10.1177/00491241251342008
[7] Sarstedt, M., Adler, S. J., Rau, L., & Schmitt, B. (2024). Using large language models to generate silicon samples in consumer and marketing research: Challenges, opportunities, and guidelines. Psychology & Marketing, 41(6), 1254–1270. https://doi.org/10.1002/mar.21982
[8] Kozlowski, A. C., & Evans, J. (2025). Simulating subjects: The promise and peril of artificial intelligence stand-ins for social agents and interactions. Sociological Methods & Research, 54(3), 1017–1073. https://doi.org/10.1177/00491241251337316
[9] Li, C., & Qi, Y. (2025). Toward accurate psychological simulations: Investigating LLMs' responses to personality and cultural variables. Computers in Human Behavior, 170, 108687. https://doi.org/10.1016/j.chb.2025.108687
[10] Zhang, J., Liang, X., Deng, A., Bonge, N., Tan, L., Zhang, L., & Zarrett, N. (2025). Leveraging interview-informed LLMs to model survey responses: Comparative insights from AI-generated and human data. arXiv preprint.
[11] Serapio-García, G., Safdari, M., Crepy, C., Sun, L., Fitz, S., Romero, P., Abdulhai, M., Faust, A., & Matarić, M. (2025). A psychometric framework for evaluating and shaping personality traits in large language models. Nature Machine Intelligence, 7, 1954–1968. https://doi.org/10.1038/s42256-025-01115-6
[12] Ferreira, G., Amidei, J., Nieto, R., & Kaltenbrunner, A. (2025). How well do simulated population samples with GPT-4 align with real ones? The case of the Eysenck Personality Questionnaire Revised-Abbreviated personality test. Health Data Science, 5, Article 0284. https://doi.org/10.34133/hds.0284
[13] Lehr, S. A., Saichandran, K. S., Harmon-Jones, E., Vitali, N., & Banaji, M. R. (2025). Kernels of selfhood: GPT-4o shows humanlike patterns of cognitive dissonance moderated by free choice. Proceedings of the National Academy of Sciences, 122(20), e2501823122. https://doi.org/10.1073/pnas.2501823122
[14] Arora, N., Chakraborty, I., & Nishimura, Y. (2025). AI–Human hybrids for marketing research: Leveraging large language models (LLMs) as collaborators. Journal of Marketing, 89(2), 1–28. https://doi.org/10.1177/00222429241276529
[15] Lim, S., Song, W., Lee, E.-J., & Jo, Y. (2025). Psychometric item validation using virtual respondents with trait-response mediators. arXiv preprint arXiv:2507.05890.
[16] Vogelsmeier, L. V. D. E., Oliveira, E., Misiejuk, K., López-Pernas, S., & Saqr, M. (2025). Delving into the psychology of machines: Exploring the structure of self-regulated learning via LLM-generated survey responses. Computers in Human Behavior, 173, Article 108769.
[17] Estevez, M., Ballestar, M. T., & Sainz, J. (2025). Market research and knowledge using generative AI: The power of large language models. Journal of Innovation & Knowledge, 10, 100796. https://doi.org/10.1016/j.jik.2025.100796
[18] Li, P., Castelo, N., Katona, Z., & Sarvary, M. (2024). Frontiers: Determining the validity of large language models for automated perceptual analysis. Marketing Science, 43(2), 254–266. https://doi.org/10.1287/mksc.2023.0454
[19] Li, Y., Liu, Y., & Yu, M. (2025). Consumer segmentation with large language models. Journal of Retailing and Consumer Services, 82, 104078.
[20] Goli, A., & Singh, A. (2024). Frontiers: Can large language models capture human preferences? Marketing Science, Articles in Advance, 1–14. https://doi.org/10.1287/mksc.2023.0306
[21] Imschloss, M., Sarstedt, M., Adler, S. J., & Cheah, J. H. (2025). Using LLMs in sensory service research: Initial insights and perspectives. The Service Industries Journal. https://doi.org/10.1080/02642069.2025.2479723
[22] Bickley, S. J., Chan, H. F., Dao, B., Torgler, B., Tran, S., & Zimbatu, A. (2025). Comparing human and synthetic data in service research: Using augmented language models to study service failures and recoveries. Journal of Services Marketing, 39(1), 36–52.
[23] Xiong, X., Wong, I. A., Huang, G. I., & Peng, Y. (2024). Understanding AI-generated experiments in tourism: Replications using GPT simulations. Journal of Travel Research, 1–17. https://doi.org/10.1177/00472875241275945
[24] Tung, Y.-H., Yang, Z.-R., Shen, M.-W., Chang, C.-Y., Chen, C.-C., & Ho, L.-C. (2025). Research note: Assessing human preferences for natural landscapes—An analysis of ChatGPT-4 and LLaVA models. Landscape and Urban Planning, 259, 105371. https://doi.org/10.1016/j.landurbplan.2025.105371
[25] Ye, Z., Yoganarasimhan, H., & Zheng, Y. (2025). LOLA: LLM-assisted online learning algorithm for content experiments. Marketing Science, Articles in Advance, 1–22. https://doi.org/10.1287/mksc.2024.0990

Downloads: 26745
Visits: 855367

Sponsors, Associates, and Links


All published work is licensed under a Creative Commons Attribution 4.0 International License.

Copyright © 2016 - 2031 Clausius Scientific Press Inc. All Rights Reserved.