Conceptio
›
Archive
›
Arxiv Oai Expanded
Arxiv Oai Expanded
metadata only
Future Policy Approximation for Offline Reinforcement Learning in LLM Reasoning
Oh, Minjae et al.
Arxiv Oai Expanded · Other
Open Source ↗
Direct PDF ↓
computation and language
This document is indexed with metadata only — full text is not available in the archive for this record.
Open the official source ↗
Related documents
Sparse PPMI Graph Averaging for Random Indexing Embeddings
#664404
Record
· ID 633297
Retrieved via
Conceptio
— every document is proof-bundled with source, license, and retrieval metadata.