Conceptio › Archive › arXiv CS
arXiv CSopen access

Intrinsic Robot Rewarding: Reusing VLA Representations for Autonomous Evaluation and Policy Improvement

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

Intrinsic Robot Rewarding: Reusing VLA Representations for Autonomous Evaluation and Policy Improvement

arXiv:2609.17115v1 [cs.RO] 15 Sep 2026

Tobias Schaffer1,* , Mohab Elkhayat1 , Daniela Nicklas1 , Mustafa Almohamad1 , Elham Al-Fuqara1

1 Technology Campus Cham - Intelligent Robotics, Deggendorf Institute of Technology, Cham, Germany * Corresponding author: [email protected]

Abstract Vision-language-action (VLA) systems already bring together two valuable resources for robot learning: rich visual representations and demonstrations of successful task execution. Intrinsic Robot Rewarding (IRR) proposes to use these resources for a second, complementary purpose: evaluating the robot’s own outcomes and providing feedback for policy improvement. Successful demonstration endpoints define task-specific references, and the policy’s frozen visual encoder provides the feature space in which new outcomes are assessed. The core reward mechanism adds a reference bank and a scoring operation to the existing pipeline, without requiring a separate learned evaluator or an additional perception backbone. Our position is that this reuse offers a promising route to lower integration effort, efficient reward computation, and reduced recurring human outcome scoring. Building on established research in visual rewards and learning from experience, IRR brings these ideas into the robot’s existing perception and demonstration pipeline. An operational COMAU Racer 3 demonstrator is available at technology readiness level 4 (TRL 4). This laboratory foundation supports the next research step: connecting internal outcome evaluation to physical policy improvement. We present the reward formulation, central research questions, and an evaluation methodology linking reward reliability to task success and supervision effort. The intended contribution is a reusable approach to learn and improve from the data and experience already available in industrial robot systems.

1 Position and motivation Vision–language–action models connect visual observations and language instructions to robot actions. RT-2 demonstrated the transfer of web-derived semantic knowledge into robotic control, while OpenVLA made a broadly pretrained VLA available for downstream adaptation [1, 2]. These developments provide an increasingly capable foundation for manipulation. Adapting such a policy to an industrial task also produces a valuable local resource with demonstrations that show both, how the task is performed and what successful execution looks like. We take the position that the representations and successful examples already available in a robot’s learning pipeline are a valuable resource for evaluating its own performance. A demonstration contains more than action supervision. Its endpoint provides evidence of the intended outcome, and the policy’s visual encoder already offers a representation in which that outcome can be described. Intrinsic Robot Rewarding puts these resources to work together: successful endpoints form a task-specific reference bank, and new attempts are scored by their relationship to those references. The central benefit of this design is reuse. In its core form, IRR adds reference management and similarity scoring to the existing perception pipeline. It requires no separate reward foundation model, no additional perception backbone, and no new reward-model training corpus for constructing the reference bank. This offers a practical opportunity to reduce integration and model-maintenance 1

Intrinsic Robot Rewarding

Research position paper

effort while making fuller use of demonstrations that have already been collected. Where reward and policy inputs coincide, visual features can also be shared at inference time. The magnitude of these benefits will be measured through computation, engineering effort, and human supervision. An intuitive analogy is learning to stack or balance objects. A learner may recognize that a stack has collapsed, or that an object has lost its balance, before being able to perform the task reliably. Recognizing the outcome and producing it consistently are related but distinct capabilities. Perceiving the consequences of an attempt provides feedback for adjusting the next one. IRR adopts this distinction at the system level: a robot’s existing visual representation may already contain information useful for recognizing successful outcomes, even while its action policy is still learning to achieve them consistently. This motivates the use of shared perceptual resources for both acting and evaluating. Reinforcement learning (RL) provides the mechanism for turning such outcome feedback into improved behavior. Geometric rewards, human outcome judgments, and dedicated vision–language model (VLM) evaluators are established ways to support this learning. IRR contributes an additional design option: obtaining the reward from resources already integrated into the task policy. This is particularly relevant when industrial systems must be adapted repeatedly to new parts, destinations, fixtures, or workspace conditions. Here, intrinsic refers to reward computation within the robot’s existing perception and learning pipeline. The task remains specified by people through instructions and successful demonstrations. IRR is therefore a task-conditioned, demonstration-derived outcome reward while curiosity approaches instead reward novelty or prediction error [3]. The intended increase in autonomy concerns repeated outcome evaluation and feedback during learning. Commissioning, independent validation, and recovery remain visible parts of the overall effort. Our available VLA demonstrator and its documented laboratory results provide a concrete foundation for this direction [4]. This paper develops the scientific position, the reward formulation, and the evidence needed to assess representation quality, physical learning, efficiency, and transfer. The reported measurements describe the existing imitation-trained system while improvements obtained through IRR are the subject of the proposed research. Together, the operational foundation and the principle of reuse establish a practical path toward self-evaluating and self-improving industrial robots.

2 Related research and proposed contribution 2.1 Policy adaptation and learning from experience OpenVLA combines DINOv2 and SigLIP visual features with a language backbone and action prediction [2, 5, 6]. Its optimized fine-tuning recipe, OpenVLA-OFT, uses parallel decoding, action chunking, continuous actions, and an L1 regression objective [7]. This provides a practical base policy for studying internal outcome evaluation. A residual RL controller can build on its regression-based action interface while supplying the stochastic learning mechanism for subsequent improvement. Progress in real robot learning gives further support to this direction. HIL-SERL integrates ∗ system demonstrations, human corrections, and efficient RL for dexterous manipulation [8]. The π0.6 and RECAP combine demonstrations, experience, and corrections in VLA learning [9]. Together, these approaches motivate the transition from imitation to learning through task experience. IRR focuses on making the associated reward provision easier to integrate by drawing on the policy’s existing perceptual representation. 2.2 Visual rewards and goal representations Value-implicit pre-training (VIP) learns a visual representation in which distance to a goal image can define rewards for downstream robot tasks [10]. LIV jointly learns language–image representations

2

Intrinsic Robot Rewarding

Research position paper

and rewards from video with text annotations [11]. RoboCLIP uses a pretrained video–language model to produce reward from a video or text demonstration [12]. Baumli et al. derive rewards for visual language goals from pretrained VLMs [13]. The listed research provides a substantial foundation for our IRR proposition that learned visual representations can support useful task feedback. IRR develops this around the resources of an existing VLA deployment. Its contribution centers on the joint use of the policy encoder and the demonstrations already collected for imitation learning, with task-specific reference construction and evaluation on industrial hardware. The relevant scientific question is how effectively this reuse supports reward generation and policy improvement, and what integration and supervision benefits it will provide. RL-VLM-F obtains preference feedback from a VLM and learns a reward function [14]. GoalLadder incrementally discovers and ranks goal states using VLM comparisons, then trains an agent to approach a selected goal in a learned embedding space [15]. These methods demonstrate complementary ways to organize goal information and feedback. IRR investigates a directly available source of such information, the successful endpoints from the task’s existing demonstrations, represented by the policy’s own visual encoder. 2.3 Self-referential rewards and recent reward models SRPO is closely related: it uses successful trajectories from the current rollout batch as references and latent representations from V-JEPA 2 to provide progress rewards for unsuccessful trajectories [16, 17]. Its formulation retains a sparse terminal success signal for identifying reference successes, and its reported physical experiments use an offline RL approach. IRR addresses a complementary question of how successful demonstration endpoints and the existing policy encoder can supply outcome feedback for repeated online physical learning. This connection places internal reference-based evaluation within an active and promising line of research. Recent reward models also provide useful methodological foundations and comparators. RoboReward introduces a real robot reward dataset, benchmark, and VLM reward models, including negative and near-miss examples [18]. Large Reward Models studies foundation VLMs adapted to generate process, completion, and temporal contrastive rewards for online refinement [19]. Their approaches inform the evaluation of outcome discrimination and learning utility. Comparisons with these models will establish where encoder reuse provides an attractive balance of task performance, computation, and supervision effort. Table 1 summarizes the research position. The bibliography reflects selected primary literature and focuses on the methods most relevant to the proposed direction.

3 Available TRL 4 demonstrator and experimental foundation An operational VLA demonstrator is available at TRL 4, providing a laboratory-validated experimental foundation for IRR. The demonstrator comprises a COMAU Racer 3 six-axis industrial arm with a Robotiq Hand-E gripper, a third-person USB camera, VR teleoperation, and a ROS control stack [4]. The baseline study uses 245 demonstrations of placing a small green cube into a container. Images, proprioceptive observations, and end-effector delta commands were recorded at a nominal 5 Hz. A seven-component action comprises six Cartesian pose increments and a gripper command. OpenVLA-7B was adapted using LoRA and the OFT recipe, the reported experiment used five-action predictions and executed two actions before querying again. Figure 1 shows a photograph of the setup. The baseline experimental evaluation consists of 50 physical trials, each with a 60-second limit, where object and container placements are varied according to a fixed protocol. The aggregate results are shown in Table 2. The largest individual failure category was failure to center the end effector on the object, accounting for 8 of the 22 failed trials. This identifies a concrete opportunity for learning from outcome feedback, particularly through corrections to approach and placement 3

Intrinsic Robot Rewarding

Research position paper

Table 1: Relationship to closely related reward approaches and the contribution investigated through IRR. Approach

Reward source or reference

IRR research distinction

VIP / LIV [10, 11]

Representations trained to support visual or language goal rewards

RoboCLIP [12]

Video/text reference and pretrained video–language encoder

RL-VLM-F / GoalLadder [14, 15] SRPO [16]

VLM feedback for reward learning or goal ranking

Can the existing policy representation suffice without reward-oriented pretraining? Can the same imitation data and policy encoder serve both action and outcome evaluation? Can task references reduce recurring external model feedback?

RoboReward / Large Reward Models [18, 19]

Successful rollouts, sparse success signal, and world-model trajectory features Separately trained VLM reward models

Can an internal visual reference supply terminal reward for online physical learning? What reliability and cost trade-off does encoder reuse achieve?

Figure 1: Photograph of the available COMAU Racer 3 laboratory cell. The setup displays the integrated Robotiq Hand-E gripper, the primary camera viewpoint, the sample cube, and the designated target container.

behavior. Comparisons with observation and demonstration improvements will establish the specific contribution of reward-driven learning. The TRL 4 demonstrator integrates the robot, teleoperation, demonstration collection, VLA adaptation, and closed-loop execution in a working laboratory system. Its 56% task success rate provides both an established baseline and a measurable opportunity for further learning. The proposed research can therefore build directly on available hardware, data, and software. Matched comparisons will use a newly measured baseline with the policy checkpoint, observations, control settings, and trial conditions held fixed.

4 Proposed IRR formulation 4.1 Task-conditioned success references Let g be a task with a written success specification, ot denote its camera observation, pt the robot’s proprioceptive state, and π0 the imitation-trained VLA policy. Let ϕ be a frozen checkpoint of its

4

Intrinsic Robot Rewarding

Research position paper

Table 2: Reported laboratory results of the available VLA demonstrator [4], providing the imitationpolicy baseline for subsequent IRR research. Measure

Reported value

Demonstration episodes Successful evaluation trials Failures due to the time limit Failures involving singularity / emergency stop Completion time among successful trials Approximate 95% Wilson interval for 28 / 50* *

245 28 / 50 (56%) 20 / 50 2 / 50 39.16 ± 9.50 s (mean ± SD) 42.3 . . . 68.8%

Calculated here from the aggregate count under a binomial model.

visual encoder and P a documented feature extraction and pooling operation. For a single image, define P (ϕ(ot )) zt = ∈ Rd , ε > 0. (1) ∥P (ϕ(ot ))∥2 + ε The encoder layer, image preprocessing, token pooling, and normalization are part of the method and must be fixed before evaluation. DINOv2 features, SigLIP features, and their fusion will be compared. A global pooled embedding is the simplest baseline. Successful endpoints from the imitation-training demonstrations define (i)

+ }. Zg+ = {zTi : τi ∈ Dg,train

(2)

Here τi is a demonstration trajectory and Ti its endpoint. The task instruction also selects the appropriate bank (a visually valid endpoint for one task must not automatically reward another). A bank represents multiple valid outcomes, such as different acceptable object poses, rather than forcing them into one visual prototype. Separate episodes and collection sessions are reserved for calibration and final testing. 4.2 Similarity as a candidate reward For 1 ≤ k ≤ |Zg+ |, let Nk (z; Zg+ ) be the k nearest reference embeddings. A simple reference distance and score are dg (z) =

1 k

∥z − u∥22 ,

X

(3)

u∈Nk (z;Zg+ )

sg (z) = exp[−dg (z)/τg ],

τg > 0.

(4)

In Equation (4), τg is the positive, task-specific distance-scale (temperature) parameter. It controls how quickly the score decreases with distance from successful reference states, smaller τg makes the score more selective, whereas larger τg gives a more gradual decay. The parameter uses the same distance scale as dg and is distinct from the demonstration trajectory τi in Equation (2). Its value is selected from training or held-out calibration data, documented for each task, and fixed before learning and final evaluation. The score lies in (0, 1] but is not a calibrated probability of success. The initial experiment will use a terminal reward, where T denotes the final time step of the attempt. To limit transient visual matches, an optional persistence score takes the minimum of the scores over a short, fixed observation window after the attempt: spers = g

min

j=0,...,m−1

sg (zT −j ),

rTIRR = qT spers g ,

rtIRR = 0 (t < T ).

(5)

The single-frame version has m = 1. The validity indicator qT rejects missing or stale observations and protocol-invalid attempts, an independent safety abort also sets reward to zero and terminates 5

Intrinsic Robot Rewarding

Research position paper

the attempt. Multiple views and ordered terminal sequences will be investigated when a single view cannot distinguish success from a near miss. 4.3 Calibration and reliable autonomous evaluation The core IRR reference model is constructed only from successful demonstrations already collected for imitation learning. For operational success detection, a task-specific threshold ηg is applied to the persistence score: ( 1, spers ≥ ηg (predicted success), g ŷ = (6) pers 0, sg < ηg (predicted failure). The threshold can be selected using held-out successful examples to obtain a desired acceptance rate. Independent failures and near misses are then used to measure the false-positive rate. This keeps reward construction based on successful demonstrations only, while using separate labeled outcomes to validate its reliability. The evaluation will compare positive-only IRR with a calibrated variant. In positive-only IRR, the reference bank and decision threshold are derived without using labeled failures; failures and near misses are reserved for independent evaluation. In the calibrated variant, a limited development set of labeled successes and failures is used either to select the decision threshold or to train a small reward head on the frozen representation. For example, threshold calibration can be performed by sweeping candidate values of ηg on the calibration set and selecting the highest threshold that still achieves the desired success sensitivity while minimizing false positives. The additional labeling effort is recorded explicitly, allowing to assess when simple reference reuse is sufficient and when modest calibration improves reliability. Figure 2 shows representative outcome images for the distinction between visual resemblance and physical task completion.

Figure 2: Representative task outcomes. From left to right: successful completion, with the object released and resting inside the target container; grasp failure where the object is not securely picked; task failure where the object is dropped and not successfully recovered.

4.4 Policy improvement and control integration The first implementation will retain the imitation-trained VLA and train a bounded residual controller for Cartesian corrections. Residual RL provides a way for combining a prior controller with learned corrections [20]. An off-policy actor–critic method such as Soft Actor-Critic (SAC) supplies a concrete stochastic learning mechanism without requiring likelihoods from the OFT regression head [21]. In schematic form, 



aexec = S a0t + Bδat , t

a0t = π0 (ot , pt , g),

δat ∼ πθ (· | zt , pt , g, a0t ).

(7)

where δat denotes the residual action correction produced by the learned policy πθ . B bounds and scales corrections in explicitly defined units and S is the fixed execution constraint layer. 6

Intrinsic Robot Rewarding

Successful demonstrations

Research position paper

encode

Camera observations

Task-specific reference bank

IRR score and validity checks

Frozen VLA visual encoder features

Industrial robot and workspace

action

VLA policy and bounded residual

update

Replay and policy improvement

imitation initialization

Figure 3: Proposed learning loop. Reference endpoints use the same frozen encoder as policy observations. The policy also receives the task instruction and proprioception. Independent outcome audits assess success outside the IRR policy-update path. Equation (7) applies to continuous Cartesian commands whereas initial gripper commands remain those of the base policy. Learning gripper timing would require an explicitly modeled discrete or hybrid action policy and is a subsequent extension. Residual improvement must consequently be reported as improvement of the combined controller, not as full VLA weight adaptation. The learner stores commanded and executed actions, interventions, timestamps, and true termination causes. If the residual policy chooses δat , replay records that choice and the resulting transition under the fixed constraint layer, it must not relabel a clipped physical command as the sampled policy action. Reward features can reuse an existing policy encoder forward pass as long as images, preprocessing, and encoder weights are identical. Full policy adaptation with a fixed reward-encoder snapshot is a later study. Likewise, dense progress shaping is an extension, because usually similarity to a final state does not increase monotonically along a valid trajectory.

5 Research questions and decisive experiments RQ1 – How effectively do existing policy representations support task evaluation? We hypothesize that a task-conditioned reference bank can distinguish successful endpoints from realistic near misses with a low false-positive rate. The evaluation will include correct placement, objects beside the container, objects still held above it, wrong-object placement, premature release, and partial occlusion. These comparisons will identify the features and reference structures that make existing representations useful for dependable outcome assessment. RQ2 – How effectively does IRR support physical policy improvement? We hypothesize that residual RL using IRR feedback improves independently measured task success over the same imitation policy and control stack. Learning curves will show both IRR scores and independently verified outcomes. Their agreement is central to demonstrating useful improvement, divergence will reveal cases for refining the evaluator. A matched human-scored RL comparison will separate the contribution of reward provision from that of the learner and control configuration. RQ3 – What efficiency benefits does reuse provide? We hypothesize that reusing the policy encoder and existing demonstrations reduces recurring human outcome-scoring time and the overhead of maintaining a separate evaluator. Evaluation will include reference construction, calibration, inference, audits, resets, and recovery. Measuring these costs together will establish the practical benefit per task and per verified successful cycle. Comparisons with dedicated reward

7

Intrinsic Robot Rewarding

Research position paper

models and further imitation training will place the benefit in the context of alternative uses of the same resources. RQ4 – How can the approach support adaptation across tasks and conditions? Task-specific reference banks offer a simple mechanism for introducing new success conditions while retaining the encoder and evaluation interface. We will study variations in lighting, viewpoint, object appearance, spatial arrangement, and task identity. Adapting to a new task through a new demonstration bank will be distinguished from zero-shot transfer. Subsequent work can extend the approach to further robot embodiments and to phase-dependent or temporal references. Where a task requires additional views, proprioception, or tactile evidence, the resulting gains and costs will be evaluated explicitly.

6 Evaluation design and evidence standards 6.1 Tasks, baselines, and ablations The first task reproduces the documented cube-to-container experiment. Two planned extensions introduce target-container placement with separation and placement into a fixture with visible pose constraints. These are laboratory tasks motivated by industrial handling. Natural-product sorting and precision assembly provide subsequent application directions, with additional acceptance measurements wherever appearance alone is inadequate. The minimum physical comparison contains: (i) the fixed imitation policy (ii) the same residual RL learner with human binary outcome rewards (iii) that learner with strict positive-only IRR and (iv) the learner with one established external visual reward baseline selected during offline screening. Candidate external rewards include VIP/LIV-style goal distance or a pretrained robot reward VLM such as RoboReward. A calibrated IRR variant can also form an additional setting. All learning conditions start from the same policy checkpoint and receive matched robot-interaction experience. The main ablations should compare a single reference with a reference bank, global with spatial features, the two visual cores (DINOv2 and SigLIP) with their fusion, a single frame with persistence, and positive-only with labeled calibration. 6.2 Independent outcomes and metrics Before data collection, each task receives an operational success definition. For pick-and-place, this includes the correct object resting inside the intended container after release for a specified observation interval, within the deadline, with no disqualifying event. Human evaluators label synchronized video without seeing the used method or reward score. Such audit labels are used to evaluate IRR, not to reward its training episodes. Table 3: Primary evaluation measures. Measure

Evaluation

Task success rate

Independently verified successes among all valid trials, report absolute differences with confidence intervals. Pr(ŷ = 1 | y = 0) on labeled failures, proposed target ≤ 2%. Pr(ŷ = 1 | y = 1) on verified successes, proposed target ≥ 90%. Verified success versus robot interactions and learning time. Time required for demonstrations, calibration, outcome scoring, audits, resets, and recovery. Completion time, retries, resets, safety interventions, and task-specific quality measures. Additional evaluator latency, memory, and computation, including extra views or frames.

False-positive rate Success sensitivity Learning efficiency Human effort Execution quality Reward cost

8

Intrinsic Robot Rewarding

Research position paper

At least 100 baseline trials will be used for diagnostic characterization. Confirmatory sample sizes depend on the expected effect and experimental replication.

7 Scientific value and industrial relevance IRR advances a practical research position, that representations and demonstrations used to establish robot behavior can also support its evaluation and improvement. Reuse can extend the function of an existing VLA deployment while keeping the reward mechanism compact. The core design introduces task references and a scoring operation, with no separate learned evaluator or additional perception backbone. This makes the approach straightforward to relate to the robot’s existing data and model configuration. The potential benefit has several dimensions. Demonstrations can support both action learning and outcome evaluation, increasing the value of the collection effort. A shared representation can reduce the number of models that must be integrated, updated, and monitored. Task references provide a direct way to introduce new success conditions, and internally computed feedback can reduce repeated human scoring during learning. These are concrete opportunities arising from the architecture, their magnitude will be established through matched performance and effort measurements. The formulation also makes it possible to study how representation quality affects reward reliability and subsequent policy improvement. Industrial handling offers a clear starting point. Variable-part manipulation, sorting into designated destinations, and placement into fixtures involve repeated adaptation to new objects and workspace conditions. For tasks with visually observable completion, demonstration-derived references offer a promising way to make reward provision easier to reuse across applications. The practical contribution will be assessed through commissioning effort, human time per verified successful cycle, throughput, and acceptance quality. The proposed IRR learning loop will operate within independently safeguarded robot operation and task-specific quality assessment. Controller limits, stopping functions, and validated recovery procedures provide the operational framework for physical learning. Where acceptance depends on forces, hidden insertion depth, or product condition, the relevant measurements complement visual outcome evaluation. This maintains a clear role for each component while allowing the reward mechanism to remain focused on learning. The direction also supports longer-term research into phase-aware progress representations, controlled expansion of reference banks, further robot embodiments, and direct VLA adaptation. Video world-model features, including V-JEPA 2 [17], provide a useful comparator for temporally structured tasks. The principle remains consistent across these extensions: make effective use of available representations and experience, and add capability where its contribution can be demonstrated.

8 Conclusion Intrinsic Robot Rewarding puts existing VLA representations and successful demonstrations to work as a source of feedback for further learning. A task-specific reference bank and compact scoring mechanism provide a concrete route to internal outcome evaluation, with the potential to reduce recurring supervision and the integration of additional reward models. The proposed work aims to connect reference-based reward generation with physical policy improvement. Its evaluation links reward reliability, verified task success, and total supervision effort, making the anticipated benefits measurable. Our position is that self-evaluating and self-improving robots can benefit substantially from making fuller use of what they already have: pretrained representations, successful demonstrations, and experience from task execution. IRR develops this principle into a research direction for

9

Intrinsic Robot Rewarding

Research position paper

industrial manipulation. It offers a promising basis for extending imitation-trained behavior toward continued learning through the robot’s own attempts.

References [1] Brianna Zitkovich, Tianhe Yu, Sichun Xu, Peng Xu, Ted Xiao, Fei Xia, Jialin Wu, Paul Wohlhart, Stefan Welker, et al. RT-2: Vision-language-action models transfer web knowledge to robotic control. In Proceedings of the 7th Conference on Robot Learning, volume 229 of Proceedings of Machine Learning Research, pages 2165–2183, 2023. URL https://proceedings.mlr.press/v229/zitkovich23a.html. [2] Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan P. Foster, Pannag R. Sanketi, Quan Vuong, Thomas Kollar, Benjamin Burchfiel, Russ Tedrake, Dorsa Sadigh, Sergey Levine, Percy Liang, and Chelsea Finn. OpenVLA: An open-source vision-language-action model. In Proceedings of the 8th Conference on Robot Learning, volume 270 of Proceedings of Machine Learning Research, pages 2679–2713, 2025. URL https://proceedings.mlr.press/ v270/kim25c.html. [3] Deepak Pathak, Pulkit Agrawal, Alexei A. Efros, and Trevor Darrell. Curiosity-driven exploration by self-supervised prediction. In Proceedings of the 34th International Conference on Machine Learning, volume 70 of Proceedings of Machine Learning Research, pages 2778–2787, 2017. URL https:// proceedings.mlr.press/v70/pathak17a.html. [4] Mohab Elkhayat, Mustafa Almohamad, and Tobias Schaffer. A vision-language-action pipeline for robotic pick-and-place: System design, OpenVLA fine-tuning, and closed-loop evaluation, 2026. Manuscript in preparation, preprint to be made available. [5] Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. DINOv2: Learning robust visual features without supervision. arXiv:2304.07193, 2023. URL https://arxiv.org/abs/2304.07193. [6] Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid loss for language image pre-training. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023. URL https://arxiv.org/abs/2303.15343. [7] Moo Jin Kim, Chelsea Finn, and Percy Liang. Fine-tuning vision-language-action models: Optimizing speed and success. In Robotics: Science and Systems, 2025. URL https://arxiv.org/abs/2502.19645. [8] Jianlan Luo, Charles Xu, Jeffrey Wu, and Sergey Levine. Precise and dexterous robotic manipulation via human-in-the-loop reinforcement learning. arXiv:2410.21845, 2024. URL https://arxiv.org/abs/ 2410.21845. ∗ [9] Physical Intelligence. π0.6 : A VLA that learns from experience. Technical report, Physical Intelligence, 2025. URL https://www.pi.website/download/pistar06.pdf.

[10] Yecheng Jason Ma, Shagun Sodhani, Dinesh Jayaraman, Osbert Bastani, Vikash Kumar, and Amy Zhang. VIP: Towards universal visual reward and representation via value-implicit pre-training. In International Conference on Learning Representations, 2023. URL https://arxiv.org/abs/2210.00030. [11] Yecheng Jason Ma, Vikash Kumar, Amy Zhang, Osbert Bastani, and Dinesh Jayaraman. LIV: Languageimage representations and rewards for robotic control. In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pages 23301–23320, 2023. URL https://proceedings.mlr.press/v202/ma23b.html. [12] Sumedh Sontakke, Jesse Zhang, Séb Arnold, Karl Pertsch, Erdem Bıyık, Dorsa Sadigh, Chelsea Finn, and Laurent Itti. RoboCLIP: One demonstration is enough to learn robot policies. In Advances in Neural Information Processing Systems, volume 36, 2023. URL https://arxiv.org/abs/2310.07899. [13] Kate Baumli, Satinder Baveja, Feryal Behbahani, Harris Chan, Gheorghe Comanici, Sebastian Flennerhag, Maxime Gazeau, Kristian Holsheimer, Dan Horgan, Michael Laskin, et al. Vision-language models as a source of rewards. arXiv:2312.09187, 2023. URL https://arxiv.org/abs/2312.09187v3. Cited version: v3, revised July 2024. [14] Yufei Wang, Zhanyi Sun, Jesse Zhang, Zhou Xian, Erdem Biyik, David Held, and Zackory Erickson. RL-VLM-F: Reinforcement learning from vision language foundation model feedback. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, pages 51484–51501, 2024. URL https://proceedings.mlr.press/v235/wang24bn.html.

10

Intrinsic Robot Rewarding

Research position paper

[15] Alexey Zakharov and Shimon Whiteson. GoalLadder: Incremental goal discovery with vision-language models. In Advances in Neural Information Processing Systems, 2025. URL https://arxiv.org/abs/2506. 16396. [16] Senyu Fei, Siyin Wang, Li Ji, Ao Li, Shiduo Zhang, Liming Liu, Jinlong Hou, Jingjing Gong, Xianzhong Zhao, and Xipeng Qiu. SRPO: Self-referential policy optimization for vision-language-action models. arXiv:2511.15605, 2025. URL https://arxiv.org/abs/2511.15605v2. Cited version: v2, 30 November 2025. [17] Mido Assran, Adrien Bardes, David Fan, et al. V-JEPA 2: Self-supervised video models enable understanding, prediction and planning. arXiv:2506.09985, 2025. URL https://arxiv.org/abs/2506.09985. [18] Tony Lee, Andrew Wagenmaker, Karl Pertsch, Percy Liang, Sergey Levine, and Chelsea Finn. RoboReward: General-purpose vision-language reward models for robotics. arXiv:2601.00675, 2026. URL https://arxiv.org/abs/2601.00675v2. Cited version: v2, 8 January 2026. [19] Yanru Wu, Weiduo Yuan, Ang Qi, Vitor Guizilini, Jiageng Mao, and Yue Wang. Large reward models: Generalizable online robot reward generation with vision-language models. arXiv:2603.16065, 2026. URL https://arxiv.org/abs/2603.16065v2. Cited version: v2, 22 March 2026. [20] Tobias Johannink, Shikhar Bahl, Ashvin Nair, Jianlan Luo, Avinash Kumar, Matthias Loskyll, Juan Aparicio Ojea, Eugen Solowjow, and Sergey Levine. Residual reinforcement learning for robot control. In IEEE International Conference on Robotics and Automation, 2019. URL https://arxiv.org/abs/1812. 03201. [21] Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In Proceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Machine Learning Research, pages 1861–1870, 2018. URL https://proceedings.mlr.press/v80/haarnoja18b.html.

11

Record · ID 919404 · SHA-256 b22784fd13d83ac2
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.