ConceptioArchivePubMed
PubMedopen access

Agreement, calibration, and failure of three large language models as high-stakes multimodal ospe graders: a comparative psychometric analysis.

Shahid Akhtar Akhund et al.
PubMed · Papers · License: Open Access
Open Source ↗
large language models, humans, educational measurement, retrospective studies, reproducibility of results, psychometrics
This document is indexed with metadata only — full text is not available in the archive for this record. Open the official source ↗

Related documents

Record · ID 358758
Conceptio Open Knowledge Archive — every document is proof-bundled with source, license, and retrieval metadata.