Calhoun Npsmetadata only
AN EMPIRICAL META-EVALUATION OF LANGUAGE MODEL EVALUATION METHODS FOR SYSTEMS ENGINEERING USING DISTRACTOR SENSITIVITY AND CONSENSUS JUDGING
systems engineering, large language models, artificial intelligence, benchmarking, evaluation metrics
This document is indexed with metadata only — full text is not available in the archive for this record.
Open the official source ↗
Related documents
Record · ID 657675
Conceptio Open Knowledge Archive — every document is proof-bundled with source, license, and retrieval metadata.