FRIDAY 11 SEPTEMBER 2026 latent·wire 108 PIECES ON FILE
← ModelsModels

Vidu S2 adds real-time 720p avatar generation and streaming video editing

The Vidu team published Vidu S2, a pair of models aimed at real-time video work, in a first version listed on arXiv cs.LG as 2609.11638v1. Vidu S2 comprises Vidu S2-Avatar, a real-time interactive digital-character model, and Vidu S2-Editing, a real-time video editing model. The same paper explores the feasibility of real-time spatial video generation for both.

Vidu S2 builds on Vidu S1, and the paper positions the avatar model as an upgrade along three lines. Vidu S2-Avatar generates 720p video in real time, takes dynamic references that can be updated at any moment, and follows instructions more strongly than its predecessor. Dancing is the example the authors give for that instruction following. The dynamic-reference capability is what separates an interactive avatar from a one-shot generator, because a character's appearance can change mid-session without restarting the run.

Vidu S2-Editing works on a video stream, changing it while it plays rather than after the fact. It applies four operations to live video: style rendering, clothing replacement, character replacement and background replacement. The abstract does not say how the stream is fed to the model or how much of the surrounding frame each edit preserves.

The spatial work appears as a feasibility question rather than a shipped capability. The authors state that they explored the feasibility of real-time spatial video generation for both Vidu S2-Avatar and Vidu S2-Editing. A playable online demo runs at vidu.com/vidu-stream, and the abstract does not say which of the two models the demo exposes.

The paper reports that experiments show Vidu S2 outperforming all baselines. It does not name the baselines, the benchmarks or the evaluation metrics, so the size of the margin and the conditions under which it holds cannot be checked against the summary. No comparison to other real-time avatar systems is offered beyond that blanket claim.

Vidu S2 arrived as a cross-listed announcement on arXiv cs.LG, with no accompanying coverage in the material available here. The abstract therefore stands as the working account of what the models do and how well they perform, and the comparison to prior work rests on the authors' own experiments. The title groups the work under three labels, interactive, editable and spatial, and the first two describe models that keep running while a user directs them rather than producing a clip and stopping.

Unknown is when Vidu S2 becomes available, at what latency, and on what hardware. The paper reports 720p generation for the avatar model but gives no resolution, frame rate or latency figure for Vidu S2-Editing, no model sizes, and no detail on training data or compute. It states no venue and no peer-review status.

Why it matters

Vidu S2 pushes the interactive-video field from generating clips toward models that run live, with avatar control and stream editing handled in one system.