Summary
During a live IsaacTeleop session with MANUS sensor collections, CloudXR’s tensor service reported XRT_ERROR_INVALID_TENSOR_ID. The OpenXR API returned XR_ERROR_RUNTIME_FAILURE, which SchemaTrackerBase treated as fatal, terminating the teleoperation session.
The immediate failure is reproducible only intermittently. The stale tensor may result from a producer disappearing or recreating its collection, but that root cause is not yet confirmed.
Environment
- IsaacTeleop version:
1.6.x
- IsaacTeleop commit:
5b5bceb7257733b6a4a7880ac5e02d3319a0de28
- CloudXR Runtime:
6.3.0
- Integration: GR00T
xr_server
- Additional device: MANUS gloves
- GR00T release: Dexter v1.2.1
Symptom
The MANUS plugin connected and registered both sensor collections:
[manus-plugin] Successfully connected to Manus host after 1 attempts
[manus-plugin] Initialized with wrist source: HandTracking
[manus-plugin] left sensors=on
[manus-plugin] right sensors=on
Shortly afterward, the next DeviceIO update failed:
ERROR [get_tensor_data] ipc_receive: Service side get_tensor_data failed:
XRT_ERROR_INVALID_TENSOR_ID
[/builds/cloudxr/cloudxr-openxr-runtime/deps/monado/src/xrt/ipc/client/ipc_client_tensor_overseer.c:111]
XR_ERROR_RUNTIME_FAILURE in xrGetTensorDataNV: Failed to get tensor data
The exception propagated through:
TeleopSession.step()
-> TeleopSession._execute_step_request()
-> DeviceIOSession.update()
-> SchemaTrackerBase::read_next_sample()
-> xrGetTensorDataNV()
xr_server then exited. The available log does not show that the OpenXR runtime process itself crashed.
Reproduction
- Start CloudXR Runtime 6.3.0.
- Start an IsaacTeleop session configured with the MANUS plugin and the
human,sensors,haptic datasets.
- Connect the headset and wait for
manus_sensors_left and manus_sensors_right to begin streaming.
- Continue teleoperation.
- Intermittently,
DeviceIOSession.update() fails with XRT_ERROR_INVALID_TENSOR_ID.
The problem has occurred alongside MANUS USB/device instability. Restarting the stack or moving the MANUS dongle and license key to different USB ports can temporarily clear it, but this is not reliable.
Expected behavior
If a tensor producer or collection temporarily disappears, the tracker should invalidate its cached collection reference and rediscover it. A transient tensor-list change should not terminate the entire teleoperation session.
Code analysis
SchemaTrackerBase::read_next_sample() handles XR_ERROR_TENSOR_LOST_NV as recoverable, but throws for every other XrResult.
The tensor extension describes XR_ERROR_TENSOR_LOST_NV as the result expected when a cached tensor-list element is no longer valid. In this case, however, the runtime’s internal XRT_ERROR_INVALID_TENSOR_ID was surfaced as the generic XR_ERROR_RUNTIME_FAILURE.
Possible explanations:
- A MANUS producer exited or recreated its tensor collection between generation checking and data retrieval.
- The CloudXR runtime retained or resolved a stale internal tensor ID.
- A producer failure occurred between IsaacTeleop’s periodic plugin-health checks, so the tensor error surfaced first.
Mitigations tested
- Restarting the complete teleoperation stack: temporarily recovers.
- Moving the MANUS USB devices to other ports: sometimes recovers, but the issue returns.
- Upgrading the GR00T deployment to Dexter v1.2.1: issue still observed.
Questions
- Why is
XRT_ERROR_INVALID_TENSOR_ID mapped to XR_ERROR_RUNTIME_FAILURE rather than XR_ERROR_TENSOR_LOST_NV?
- Should
SchemaTrackerBase invalidate and refresh its tensor list after this failure?
- Would one forced refresh and retry allow recovery while preserving fatal handling for persistent runtime failures?
- Can the runtime log the missing tensor ID, current generation, and owning producer when this occurs?
Additional context
The same investigation also observed separate MANUS plugin processes terminating with SIGABRT and SIGSEGV. Those failures may explain a disappearing tensor producer, but they are being tracked separately and should not be assumed to share this root cause.
Full station logs are available. Sensitive MANUS license information has been omitted.
Summary
During a live IsaacTeleop session with MANUS sensor collections, CloudXR’s tensor service reported
XRT_ERROR_INVALID_TENSOR_ID. The OpenXR API returnedXR_ERROR_RUNTIME_FAILURE, whichSchemaTrackerBasetreated as fatal, terminating the teleoperation session.The immediate failure is reproducible only intermittently. The stale tensor may result from a producer disappearing or recreating its collection, but that root cause is not yet confirmed.
Environment
1.6.x5b5bceb7257733b6a4a7880ac5e02d3319a0de286.3.0xr_serverSymptom
The MANUS plugin connected and registered both sensor collections:
Shortly afterward, the next DeviceIO update failed:
The exception propagated through:
xr_serverthen exited. The available log does not show that the OpenXR runtime process itself crashed.Reproduction
human,sensors,hapticdatasets.manus_sensors_leftandmanus_sensors_rightto begin streaming.DeviceIOSession.update()fails withXRT_ERROR_INVALID_TENSOR_ID.The problem has occurred alongside MANUS USB/device instability. Restarting the stack or moving the MANUS dongle and license key to different USB ports can temporarily clear it, but this is not reliable.
Expected behavior
If a tensor producer or collection temporarily disappears, the tracker should invalidate its cached collection reference and rediscover it. A transient tensor-list change should not terminate the entire teleoperation session.
Code analysis
SchemaTrackerBase::read_next_sample()handlesXR_ERROR_TENSOR_LOST_NVas recoverable, but throws for every otherXrResult.The tensor extension describes
XR_ERROR_TENSOR_LOST_NVas the result expected when a cached tensor-list element is no longer valid. In this case, however, the runtime’s internalXRT_ERROR_INVALID_TENSOR_IDwas surfaced as the genericXR_ERROR_RUNTIME_FAILURE.Possible explanations:
Mitigations tested
Questions
XRT_ERROR_INVALID_TENSOR_IDmapped toXR_ERROR_RUNTIME_FAILURErather thanXR_ERROR_TENSOR_LOST_NV?SchemaTrackerBaseinvalidate and refresh its tensor list after this failure?Additional context
The same investigation also observed separate MANUS plugin processes terminating with
SIGABRTandSIGSEGV. Those failures may explain a disappearing tensor producer, but they are being tracked separately and should not be assumed to share this root cause.Full station logs are available. Sensitive MANUS license information has been omitted.