ConceptioArchiveKaist Koasas
Kaist Koasasmetadata only

Towards Practical and Efficient Image-to-Speech Captioning with Vision-Language Pre-training and Multi-modal Tokens

Kim, Minsu et al.
Kaist Koasas · Other
Open Source ↗
academic, research, south korea
This document is indexed with metadata only — full text is not available in the archive for this record. Open the official source ↗

Related documents

Record · ID 436426
Conceptio Open Knowledge Archive — every document is proof-bundled with source, license, and retrieval metadata.