A dynamic-world representation for Physical AI

Physical AI is not only about generating pixels; it is about sensing, reconstructing, and querying a changing real world. Conventional 4DGS commonly starts from synchronized, calibrated multi-view video. This article explores a lower-capture-barrier path: start from monocular video, synthesize synchronized multi-view observations with a 4D video-generation model, and build a queryable 4DGS representation of the dynamic world.

Within rocPAI-Forge’s Physical AI direction, this pipeline is a dynamic visual layer:

monocular video → generative multi-view → 4DGS dynamic-world representation
                 → robot observations / simulation / policy training → real-robot validation

4DGS is neither the final controller nor a complete physics simulator. It represents how the world appears to change in space and time; collision, friction, contact dynamics, and joint control still belong to systems such as MuJoCo, Isaac Sim, or other robotics and physics stacks.

What we built

Turning a character video into a dynamic scene that can be viewed from changing cameras takes more than a video-generation model. Conventional 4DGS commonly starts from synchronized, calibrated multi-view video. Phi Media Lab + rocPAI-Lab explored a lower-capture-barrier path: start from a monocular video, use a 4D video-generation model on AMD Instinct MI300X + ROCm to synthesize synchronized multi-view observations, reconstruct 4D Gaussians, and then use an AMD Radeon node for Vulkan rendering, VA-API H.264 encoding, and WebRTC delivery.

This does not treat generated views as equivalent to real cameras. It trades lower capture, synchronization, and calibration overhead for additional generation compute and multi-view consistency validation.

One end-to-end run produced 24 views and 2,904 704×1280 observations from a 121-frame monocular character-video interval, initialized 499,972 dynamic Gaussians, ran 30,000 updates, exported an AssetBundle, and delivered 1280×720 motion to Chrome over WebRTC.

This is a result from one validated software and hardware path, not a cross-platform benchmark.

4DGS dynamic preview rendered by the AMD Radeon reference Viewer

Dynamic rendering preview: the AMD Radeon Viewer renders 4DGS while camera trajectory and normalized time change. The media comes from 4DGS Viewer Phi; only the rendered result is distributed, not the source video or Gaussian asset.

Why a reusable asset matters

A fixed video delivers known pixels. Free-viewpoint browsing and repeated camera design query the same content with changing camera and time parameters. A 4DGS asset makes that content reusable:

generate every frame: C_gen(N)   = gN
build an asset first:  C_asset(N) = B + rN

B is the generation and reconstruction cost; r is the cost of one rendered query. The asset path is not automatically cheaper, but it becomes attractive when the content is queried repeatedly and g > r.

Three projects, one pipeline

Generation accepts a monocular character video and produces synchronized multi-view observations. Reconstruction optimizes a continuous-time Gaussian representation. Viewer consumes the exported AssetBundle and uses Rust, wgpu/WGSL, Mesa RADV/Vulkan, and VA-API/GStreamer to deliver a WebRTC stream.

The MI300X/ROCm side builds the asset; the Radeon viewer does not need the training environment. GPU sharing avoids an unnecessary full-frame CPU round trip, while network transport and browser decoding remain part of the end-to-end path.

System architecture across Generation, Reconstruction, Viewer, MI300X, and Radeon

From one pipeline to interactive dynamic content

MP4 packages and plays already-determined pixels. 4DGS explores a different delivery boundary: a dynamic scene asset queried by camera and time. If asset schemas, compression, cross-platform rendering, and client support mature, 4DGS could become a foundation for interactive dynamic content. It would not simply replace MP4, which remains strong for fixed playback and low-bandwidth distribution.

For content creators, one capture and reconstruction could serve multiple camera paths, free-viewpoint browsing, XR, digital humans, and product presentation. For platforms, the asset becomes a basis for hosting, versioning, cloud rendering, view queries, and WebRTC delivery instead of regenerating a complete video for every new shot. The user moves from watching a fixed edit to exploring time-varying content.

For robotics, the more precise role is a dynamic visual layer, not a complete simulator. A 4DGS asset can provide queryable observations for camera changes, action replay, occlusion inspection, visual localization, and policy-data augmentation. It does not independently provide collision handling, friction, contact dynamics, joint control, or reliable prediction of unobserved actions:

real video → 4DGS visual asset → robot observations / data generation
           → physics simulation and policy training → real-robot validation

Moving from a research representation to an industry format also requires a stable asset schema, compression and streaming, cross-GPU rendering, quality validation, provenance and rights tracking, and a client ecosystem. In the near term, 4DGS is more likely to land as an application asset, cloud-rendering format, or robotics data intermediate than as a direct replacement for MP4.

Lessons and limits

We used ROCm PyTorch and SDPA for generation, locked reconstruction to gfx942, and used the AMD-Ecosystem gsplat path. When one FP16 foreground-segmentation path triggered a rocBLAS failure, we changed that branch to FP32 instead of forcing one precision across the whole pipeline.

The run demonstrates a connected AMD path from generation to interactive delivery. It does not solve hair, fast non-rigid motion, unseen-view blur, or multi-tenant serving. The next article will focus on dynamic Gaussian initialization, differentiable rendering, and separating geometry validation from performance measurement.

Project repositories