Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,6 +25,12 @@ runtime path, supported workloads, and reproducible real-time gate.

## News 📰

- ✨ **2026-09-14**: Established a general VLA session-serving abstraction with validated semantic observation and
action contracts, extensible embodiment adapters, replica-affine sessions, continuous-control WebSocket serving,
and simulator-side MuJoCo validation.
- ✨ **2026-08-27**: Added [**LingBot-VLA v2**](examples/lingbot_vla_v2/README.md) 6B base-checkpoint support
with RobotWin multi-camera, task-text, robot-state, and normalized action I/O, together with native structured HTTP
serving, BF16 eager, CUDA Graph, and quantized inference support.
- ✨ **2026-08-19**: Added [**LTX-2.5 Distilled**](examples/ltx25_distilled/README.md) T2V and I2V joint
audio-video generation with a ModuleManager-backed six-stage pipeline, selectable dense attention backends, and
Ulysses sequence parallelism on **1, 2, or 4 x H100** GPUs.
Expand Down
4 changes: 4 additions & 0 deletions README_zh.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,6 +24,10 @@ TeleFuser 是一个开源的多模态生成与世界模型流式推理和服务

## News 📰

- ✨ **2026-09-14**:建立通用 VLA 会话服务抽象,定义可校验的语义观测与动作契约,
支持可扩展具身适配器、模型副本绑定会话、连续控制 WebSocket 服务及模拟器侧 MuJoCo 验证。
- ✨ **2026-08-27**:接入 [**LingBot-VLA v2**](examples/lingbot_vla_v2/README.md) 6B base checkpoint,支持 RobotWin 多相机、
任务文本、机器人状态和归一化动作输入输出,并提供原生结构化 HTTP、BF16 eager、CUDA Graph 及量化推理支持。
- ✨ **2026-08-19**:新增 [**LTX-2.5 Distilled**](examples/ltx25_distilled/README.md) T2V 和 I2V 联合
音视频生成,采用基于 ModuleManager 的六阶段 Pipeline,支持选择密集注意力后端,并可在
**1、2 或 4 张 H100** 上使用 Ulysses 序列并行。
Expand Down
62 changes: 40 additions & 22 deletions docs/en/vla.md
Original file line number Diff line number Diff line change
Expand Up @@ -34,13 +34,14 @@ The public modules are:
- `telefuser.vla.session`: transport-neutral OPEN, PREDICT, RESET, and CLOSE lifecycle.
- `telefuser.vla.serialization`: versioned JSON/Base64 wire formats for action and observation spaces, observations,
and chunks.
- `telefuser.vla.runtime`: scheduling, deterministic chunk state, action trimming, age checks, and safety policies.
- `telefuser.vla.runtime`: deterministic chunk state, action trimming, age checks, safety policies, and simulator
execution reports. Session admission and latest-wins request handling live in `telefuser.service.vla_session`.
- `telefuser.integrations.sim`: simulator protocol and the dependency-free RoboTwin callback adapter.
- `telefuser.service.vla_session`: the additive generic VLA WebSocket application factory.
- `telefuser.service.vla_replica`: worker-local OPEN, PREDICT, RESET, and CLOSE dispatch for pipeline replicas.
- `telefuser.client.AsyncVLAClient`: remote session client with concurrent request correlation.

## LingBot-VLA v2 Compatibility
## LingBot-VLA v2 Integration

`LingBotVlaV2VLAPolicy` wraps `LingBotVlaV2Pipeline`; the pipeline's existing tensor input and return types are not
changed. It labels the normalized canonical `[T,55]` result as a `ModelActionChunk`.
Expand All @@ -49,16 +50,14 @@ changed. It labels the normalized canonical `[T,55]` result as a `ModelActionChu
chunk to an absolute-position `[H,14]` `RobotActionChunk` in the declared dual-arm joint order. Its model and robot
action spaces are exported as `LINGBOT_VLA_V2_ACTION_SPACE` and `ROBOTWIN_ACTION_SPACE`.

The generic VLA WebSocket is the primary online integration path. The existing LingBot RoboTwin WebSocket endpoint
is a compatibility adapter for unmodified upstream `WebsocketClientPolicy` clients and calls the same policy,
embodiment, session, and runtime path internally. Its URL, MessagePack request fields, metadata frame, response fields,
reset behavior, and latest-wins
scheduler behavior remain compatible. The old model-specific scheduler import is retained as an alias to the common
runtime scheduler.
The generic VLA WebSocket is the online integration path for both RoboTwin and MuJoCo clients. Each connection opens
an explicit model and embodiment pair, while the shared HTTP structured-task route remains unchanged and continues to
return the existing canonical JSON result.

The standalone endpoint owns one resident pipeline, so each connection session is inherently pinned to that policy
instance. The shared HTTP structured-task route remains unchanged and continues to return the existing canonical
JSON result.
LingBot-VLA v2 exposes three access modes over the same pipeline: `lingbot_vla_v2_inference.py` is the direct offline
baseline, `lingbot_vla_v2_native_service.py` is the existing structured HTTP compatibility path, and
`lingbot_vla_v2_vla_server.py` is the continuous-control WebSocket path. Only the WebSocket path returns the decoded
semantic `RobotActionChunk`; the direct and native paths retain the canonical normalized model result.

## Session Lifecycle

Expand Down Expand Up @@ -102,9 +101,7 @@ robot action and observation spaces; semantic incompatibility is rejected before
The protocol returns machine-readable error codes for malformed messages, unsupported versions, unknown components
or sessions, contract mismatches, superseded work, out-of-order or expired observations, request timeout, unavailable
sessions, and replica failure. Tensor payloads are dense, typed, shape-checked Base64 data inside bounded JSON
messages. This is the stable interoperability format; the LingBot-RoboTwin compatibility endpoint continues to use
its existing MessagePack format. New model and simulator integrations must target the generic protocol rather than
add behavior to the model-specific compatibility endpoint.
messages. This is the stable interoperability format for all model and simulator integrations.

`request_ttl_ms` is measured with server monotonic time and bounds inference delivery. Observation age is checked
separately using `observation_timestamp_ns` and `observation_clock_now_ns`, which must come from the same clock domain.
Expand Down Expand Up @@ -140,9 +137,34 @@ claim that a returned action was executed. `SimulatorChunkRuntime` runs on the s
EXECUTING, EXECUTED, HOLD, and STOP transitions while calling a `SimulatorAdapter` one action at a time.

The bundled RoboTwin statistics do not declare an authoritative simulator control frequency, so both exported
LingBot/RoboTwin specs use `control_hz=None`. The remote simulator must resolve its actual control period rather than
LingBot/RobotWin specs use `control_hz=None`. The remote simulator must resolve its actual control period rather than
guessing one in the inference process.

## LingBot-VLA v2 Verification Scope

The TeleFuser integration consumes the official checkpoint as-is. It does not add post-training, fine-tuning, or
model-specific camera-pose inputs. The supported verification path is:

```text
direct pipeline baseline
-> native structured HTTP compatibility
-> generic VLA WebSocket session
-> simulator-side RobotWin execution
```

The direct and native paths return the canonical normalized `50 x 55` action chunk. The generic session path applies
`RobotWinProfile` and returns an absolute-position `50 x 14` `RobotActionChunk`. MuJoCo exercises the simulator-side
execution boundary; it does not change model computation or action semantics.

Keep inference speed comparisons and release artifacts separate from task-success claims. The maintained validation
tools cover upstream tensor parity, runtime latency, quantization comparisons, structured-service behavior, replica
faults, and shutdown/restart. Generated captures and reports belong under the ignored `work_dirs/` directory. A
real-robot task-success evaluation is outside this integration and must not be represented by `policy_verified`.

The generic WebSocket remains an explicit standalone service so existing `telefuser serve` routes and non-VLA
pipelines stay unchanged. Deployment authentication, TLS, actuator feedback, and emergency-stop behavior remain
outside the inference process and must be supplied by the deployment or simulator boundary.

## Simulator Boundary

`RoboTwinSimulatorAdapter` has no RoboTwin, SAPIEN, Vulkan, or ROS dependency. The RTX-side process provides observe,
Expand All @@ -160,14 +182,10 @@ route. The LingBot example supplies a standalone server that starts `PipelinePoo
This keeps existing HTTP routing and every non-VLA pipeline unchanged. A real simulator still owns control timing,
actuator feedback, and verification that each returned action was actually applied.

For LingBot-VLA v2, the standalone generic WebSocket is the primary continuous-control entrypoint. The direct Python
For LingBot-VLA v2, the standalone generic WebSocket is the continuous-control entrypoint. The direct Python
entrypoint remains the reference/offline baseline, and the native HTTP structured service remains for existing
TeleFuser callers. The RoboTwin MessagePack server is a legacy compatibility entrypoint only; it can be removed after
all upstream clients migrate to `/v1/vla/session`.

The direct inference CLI and structured HTTP task remain separate because they provide reference and batch workflows,
not simulator session transports. The legacy RoboTwin MessagePack endpoint should be removed only after the RTX client
passes generic-protocol action delivery, reset, timeout, reconnect, and latest-wins parity checks.
TeleFuser callers. The direct inference CLI and structured HTTP task remain separate because they provide reference
and batch workflows, not simulator session transports.

No new dependencies, environment variables, shared model configuration fields, CLI options, HTTP schemas, or
existing service routes are introduced by this semantic and transport layer.
Loading
Loading