arXiv (All)open access
MUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Education
artificial intelligence, computation and language, computer vision and pattern recognition
This document is indexed with metadata only — full text is not available in the archive for this record.
Open the official source ↗
Related documents
Record · ID 1013296
Retrieved via
Conceptio — every document is proof-bundled with source, license, and retrieval metadata.