Skip to content

fix(appo): keep Motrix playback observations in tensors - #1731

Merged
TATP-233 merged 2 commits into
develop/tensor-runtimefrom
feat/issue-1704-tensor-appo-play-loop
Sep 29, 2026
Merged

TATP-233 merged 2 commits into
develop/tensor-runtimefrom
feat/issue-1704-tensor-appo-play-loop

Conversation

@TATP-233

Copy link
Copy Markdown
Collaborator

Reviewable outcome

The APPO/Motrix playback loop now keeps policy I/O on the strict TorchEnv tensor boundary:

  • reset indices are allocated on env.device as int64 tensors;
  • actor observations stay float32 Torch tensors and fail closed if a non-tensor reaches the reset or step boundary;
  • TensorDict policy input is moved explicitly to the actor device;
  • actor output is moved directly to authoritative env.device as contiguous float32 and submitted to TorchEnv.step();
  • the next actor observation is extracted from the returned TorchEnvState without NumPy or a host roundtrip.

Closes part of #1704.

Validation

  • UNILAB_LOCAL_UNISIM=/home/user/ws/unilabsim/unisim uv run pytest tests/scripts/test_train_scripts.py -q -k 'run_motrix_play_loop' --override-ini='markers=' -m '' — 1 passed.
  • UNILAB_LOCAL_UNISIM=/home/user/ws/unilabsim/unisim make check — pass (mypy clean; one pre-existing unresolved optional drake_uni.runtime pyright warning).
  • UNILAB_LOCAL_UNISIM=/home/user/ws/unilabsim/unisim make test-all — 1834 passed / 26 skipped, benchmark imports 35/35 and 36/36.
  • Full slow selection for tests/scripts/test_train_scripts.py has three pre-existing failures on the unchanged base (two off-policy video return assertions and one play_interactive physics-state contract failure); make test-all deselects that slow lane and the focused changed test passes.

Notes for review

Independent of #1725–#1730 and based on current develop/tensor-runtime.

@TATP-233
TATP-233 requested a review from caozx1110 as a code owner September 29, 2026 11:07
@TATP-233
TATP-233 merged commit 5bbe4bd into develop/tensor-runtime Sep 29, 2026
2 checks passed
@TATP-233
TATP-233 deleted the feat/issue-1704-tensor-appo-play-loop branch September 29, 2026 13:08
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant