Back to search

Article

Evaluating Large Language Models for Psychometric Simulation Studies in R: Integrating Best Practices in Simulation and Prompt Design

2026-04-08

Abstract excerpt

<p>While large language model (LLM) capabilities have been widely investigated across various domains, limited research has examined their capability to generate end-to-end R code for conducting simulation studies in psychometrics. To address this gap, we evaluated four commonly used reasoning models: ChatGPT-5.1 (Thinking), Claude 4.5 Opus, DeepSeek-V3.2, and Gemini 3 Pro, in generating simulation code correspond...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
db68a766-73d2-51f8-8732-5179c97c7b3b
DOI
10.31234/osf.io/9tfej_v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Evaluating Large Language Models for Psychometric Simulation Studies in R: Integrating Best Practices in Simulation and Prompt DesignDOI 10.31234/osf.io/9tfej_v1
Select a neighboring publication to make it the new centre.