JEPA-TTT: Persistent Test-Time Training of Latent World Models for Planning under Dynamics Shifts

1Honda Research Institute USA 2Johns Hopkins University

†Work done during an internship at Honda Research Institute USA.

World Models in Physical AI Workshop @ NeurIPS 2026

JEPA-TTT loop: encode observations with a frozen encoder, plan with the trainable predictor and frozen reward head, execute actions, and update the predictor using dense replay. Updates persist across episodes.
A world model that keeps learning as it acts. JEPA-TTT updates its latent dynamics predictor from observed transitions and retains adaptation across episodes. The visual encoder and reward head stay fixed; planning needs neither goal images nor online environment rewards.

Abstract

World models enable agents to plan by predicting future states of the environment, but their predictions can become unreliable when test-time dynamics differ from those seen during training. We present JEPA-TTT, which adapts the latent dynamics predictor of a pretrained action-conditioned Joint-Embedding Predictive Architecture world model throughout test time. Self-supervised updates accumulate across episodes, while the visual encoder and reward head remain fixed, preserving the pretrained representation and task objective. Planning requires neither a goal image nor online environment reward. JEPA-TTT uses dense replay, which forms prediction windows at every temporal offset, retains them in a growing buffer, and samples minibatches from that buffer for predictor updates. Across eight dynamics shifts in four continuous-control environments, JEPA-TTT improves planning on every shift. After 500 test-time episodes, it reduces autoregressive latent prediction error by 83% on average and improves planning performance by 153% over the frozen JEPA world model. These results show that persistent self-supervised test-time training can adapt a pretrained latent world model under changed dynamics.

Persistent Test-Time Training

When dynamics change, the same action can produce a different outcome. A frozen world model can then rank candidate plans incorrectly. JEPA-TTT uses the agent’s own observations and executed actions to adapt the predictor throughout deployment.

A candidate action has different outcomes under training and test dynamics, causing a pretrained world model to make an incorrect prediction.
A dynamics shift changes the consequences of an action. The observed transition provides a self-supervised target for adapting the dynamics model.
  1. Plan and act. Cross-Entropy Method (CEM) planning rolls candidate actions forward in latent space and scores them with an offline-trained, frozen reward head. Execute the first action block, observe the result, and replan.
  2. Learn from dense replay. Form prediction windows at every valid temporal offset and store them in a growing buffer. Sample minibatches and update the predictor to match the frozen encoder’s representations of observed states.
  3. Carry adaptation forward. Retain the predictor parameters, optimizer state, and replay buffer across episode boundaries. At most one predictor update is performed per replanning step when a full minibatch is available.

Only the dynamics predictor is adapted. Keeping the visual encoder and reward head frozen preserves both the representation space and the task objective. Reward labels are used to train the reward head offline; environment rewards are not observed during test-time training or planning.

Eight Dynamics Shifts, Four Environments

We evaluate two controlled dynamics shifts in each of PushT, Two-Room, Reacher, and OGBench-Cube. Each deployment has one dynamics change that remains fixed for 500 test-time episodes. The observation mapping, task score, action space, and episode horizon stay unchanged.

Eight shifts: PushT action and contact rotations; Two-Room spatial wave and grid rotation fields; Reacher joint-phase and harmonic fields; OGBench-Cube quadratic crosstalk and delayed cyclic actuation.
Gray dashed arrows indicate training dynamics; orange arrows indicate test dynamics. The shifts span actuator mappings, contact response, state-dependent control, cross-axis coupling, and delayed actuation.
What changes at test time?
EnvironmentDynamics shifts
PushTRotate the agent’s action by 90°, or rotate only the block’s contact-induced motion by 90°.
Two-RoomRotate actions according to a spatial wave field or a two-dimensional grid field.
ReacherRotate motor commands using a joint-phase field or a harmonic field that depends on joint configuration.
OGBench-CubeAdd nonlinear cross-axis coupling, or delay the command by one step and route it cyclically across axes.

Each world model is pretrained on 3,000 episodes under the training dynamics. During deployment, we evaluate every 50 episodes on the same 100 held-out episodes, with three deployment runs per shift. Held-out evaluation trajectories never enter the adaptation buffer.

Planning Improves on Every Shift

JEPA-TTT outperforms Frozen JEPA, PPO-TTT, and the adapted AdaJEPA baseline on all eight shifts for both reported planning metrics. Averaged across shifts, the best held-out score improves by 153% and the normalized planning-score AUC improves by 113% over Frozen JEPA.

Best held-out planning score and normalized score AUC for all eight shifts. JEPA-TTT has the highest score in every comparison against Frozen JEPA, PPO-TTT, and AdaJEPA.
Planning performance under shifted dynamics. Panel (a) shows best held-out scores; panel (b) shows normalized held-out score AUC over episodes 0–500. Error bars report standard deviations across three deployment runs.
Planning results averaged across the eight shifts. Higher is better.
MethodBest scoreMean AUCOnline environment rewards
Frozen JEPA0.2670.267No
PPO-TTT0.3070.279Yes (oracle rewards)
AdaJEPA0.2970.297No
JEPA-TTT0.6780.571No

Reading the metrics. Best score is the maximum three-run mean over evaluations at episodes 50, 100, …, 500. Mean AUC is the area under the held-out score curve from episode 0 to 500, divided by 500 and averaged over runs. It measures performance throughout adaptation.

Controlled comparisons. Frozen JEPA and AdaJEPA use the same pretrained world model, frozen reward head, and CEM planner as JEPA-TTT. AdaJEPA is adapted to this goal-image-free setting and resets its adaptation each episode. PPO-TTT receives online environment rewards during its 500-episode adaptation.

View planning curves over all 500 episodes
Eight learning curves compare JEPA-TTT with Frozen JEPA, PPO-TTT, AdaJEPA, and an offline reference across 500 test-time episodes.
Curves show means across three deployment runs; shaded regions show standard deviations. On six of eight shifts, JEPA-TTT’s best score is close to or slightly exceeds the data-rich offline reference. That reference uses 3,000 episodes collected directly under the test dynamics and is neither a deployable baseline nor a mathematical upper bound.

More Accurate Dynamics, Better Plans

After 500 test-time episodes, JEPA-TTT reduces five-block autoregressive latent prediction MSE by 83% on average. Error decreases on every shift when the frozen and adapted models receive identical held-out trajectories, observed context, and future actions.

Autoregressive latent MSE decreases after adaptation on all eight shifts, with relative reductions ranging from approximately 45% to 99%.
The predictor rolls out five future latent states without intermediate observations. Percentages indicate relative MSE reduction after adaptation; error bars show standard deviations over three adapted models. Raw MSE values are not directly comparable across tasks because their latent spaces are trained independently.
Qualitative examples show Frozen JEPA incorrectly favoring Plan A while JEPA-TTT correctly favors Plan B, which achieves a higher ground-truth task score.
Correcting the ordering of candidate plans. Plan A is optimized with Frozen JEPA and Plan B with JEPA-TTT from the same initial state. The adapted predictor recovers the better plan’s ranking while using the same frozen reward head. A pretrained decoder is used only to visualize predictions.
View decoded predictions before and after adaptation
Ground-truth, frozen-model, and adapted-model image sequences for both PushT and Two-Room shifts at 5, 15, and 25 steps ahead.
PushT and Two-Room: ground truth, Frozen JEPA, and the adapted model share the same initial observations and future action sequence. Columns show 5, 15, and 25 raw environment steps ahead.
Ground-truth, frozen-model, and adapted-model image sequences for both Reacher and OGBench-Cube shifts at 5, 15, and 25 steps ahead.
Reacher and OGBench-Cube, with the same layout. These decoded rollouts are qualitative diagnostics; neither the decoder nor its output is used for planning or computing latent prediction MSE.

What Makes Adaptation Effective?

Controlled ablations separate additional optimization, temporal coverage, and replay. Dense temporal windows improve aggregate planning performance, with a small additional benefit from replay.

Mean across eight shifts, with standard deviations over three run-level averages.
Update ruleBest scoreMean AUC
Sparse stream0.497 ± 0.0280.368 ± 0.007
Compute-matched sparse stream0.625 ± 0.0250.495 ± 0.025
Dense stream0.671 ± 0.0080.548 ± 0.007
Dense replay (JEPA-TTT)0.678 ± 0.0130.571 ± 0.004

Sparse windows begin at action-block boundaries; dense windows begin at every environment step. Compute-matched sparse streaming repeats sparse windows to match the dense presentation budget. Dense replay samples from the accumulated dense windows under the same cumulative update budget as the dense stream.

Persistence matters. In a separate matched control using episode-local dense windows, retaining the predictor and optimizer state across episodes raises best score from 0.289 to 0.729 and mean AUC from 0.273 to 0.637, with the same number of predictor updates. This control uses a batch size of 8 and is separate from the main dense replay protocol.

Scope. The study evaluates a single persistent dynamics shift. Simultaneous visual or reward changes, sequences of dynamics changes, and forgetting of previously learned dynamics remain outside the evaluation.

Citation

If you find this work useful, please cite it using the BibTeX entry below.

@article{zhang2026jepattt,
  title     = {{JEPA-TTT}: Persistent Test-Time Training of Latent World Models
               for Planning under Dynamics Shifts},
  author    = {Zhang, Zheyuan and Ye, Suyu and Agarwal, Nakul and
               Mahjoub, Hossein Nourkhiz and Pari, Ehsan Moradi and
               Khashabi, Daniel and Shu, Tianmin and Tadiparthi, Vaishnav},
  journal   = {NeurIPS 2026 Workshop on World Models in Physical AI},
  year      = {2026},
  url       = {https://www.alphaxiv.org/abs/2609.jepa-ttt}
}

Download citation