Real-world livestream intelligence · Preprint

Live Assistant: Learning Whether, When, and Whom to
Assist in Real-World Live Social Streams

Selective proactive assistance as a decision trajectory

Live Assistant learns whether to act, when assistance is needed, and whom to address — turning passive video understanding into timely participation.

Shujian Gao1,2,3Jiamei Yan2Yuchen Yang2 Penghao Zhou2,†Qinglei Wang2,†Tiehan Fan2 Yuan Wang4Zuxuan Wu1,3,*Yu-Gang Jiang1,*
1 Fudan University2 ByteDance TikTok3 Shanghai Innovation Institution4 Zhejiang University
* Corresponding author  ·  † Project lead
The problem, at a glance

From answering queries to discovering needs

Instead of waiting for a prompt, LiveAssistant continuously reads the social stream and chooses whether to observe, remember, or assist the right participant.

Comparison between query-driven video answering and LiveAssistant's proactive decision trajectory
LiveAssistant evaluates a causal decision trajectory: whether to act, when to act, whom to address, and what to say.
Live qualitative cases

Watch the policy think in time

Each frame is one step in a continuous trajectory. The model can wait, preserve context, or produce a role-aware response only when the stream calls for it.

LIVE TRAJECTORY
10-minute benchmark overview

One stream. Sixty chunks. One selective policy.

This complete 10-minute example visualizes all 60 causal decisions, from state selection to recipient routing, and explains why useful assistance sometimes means not emitting anything at all.

Ten-minute benchmark case with 60 ten-second chunks and worked decisions
A full 60 × 10-second trajectory: 19 benchmark emits and 41 no-emits, with worked decisions for host assistance, viewer assistance, and deliberate restraint.
Core idea

Silence is not a failure. It is a decision.

Most streaming models are optimized to keep talking. LiveAssistant reframes livestream understanding as selective participation in a shared, multi-party environment.

01<obs>

Observe

Keep watching without disturbing the stream when no intervention is needed.

02<mem>

Remember

Write a private clue to memory so future decisions retain causal context.

03<ans>

Assist

Speak to viewers, hosts, or moderators with a specific, grounded task.

VideoAudioCommentsGiftsViewer dynamicsRoom metadata
Why LiveAssistant

Built for the social dynamics of live video

A native livestream task requires more than frame understanding. It must reason causally over time, route actions across roles, and learn restraint.

01

A new problem, not another score

Causal, mixed-initiative decision-making over heterogeneous, long-horizon signals.

02

Multi-party routing

Viewer, host, and moderator are first-class policy decisions.

03

Omni-native input

Video, audio, comments, gifts, dynamics, and metadata are fused directly.

04

Trajectory data engine

Aligned 10-second decisions with automatic checks and human review.

05

Two-stage training

MA-MSFT learns the grammar; SM-GSPO optimizes the streaming policy.

Method

One continuous path from signals to actions

The same decision trajectory connects annotation, supervised fine-tuning, reinforcement learning, and bounded-cache inference.

LiveAssistant data, training, and inference pipeline
Data and training pipeline: causal multimodal trajectories, marker-aware SFT, and streaming-oriented policy optimization.
01

Serialize live-native signals

Align video, audio, interaction, gifts, dynamics, and metadata in causal order.

02

Learn the action grammar

MA-MSFT protects rare structural decisions with marker-aware objectives.

03

Optimize trajectories

SM-GSPO combines structure, content, turn-level, and trajectory-level rewards.

04

Reason over long streams

A dense recent window and compressed long-term memory preserve useful context.

Benchmark & results

Strictly causal. Human verified. Room-disjoint.

Evaluation uses native audio, video, comments, and gifts. The non-trivial state distribution penalizes both “always talk” and “always silent” policies.

38.17hbenchmark video
13,812unique chunks
625Kcomments
9content categories
State distribution
OBS 48.2%MEM 25.9%ANS 25.9%

Controlled ablation Progressive training strategy, %

Training strategyState
overall
OBS
recall
MEM
recall
ANS
recall
Recipient
overall
Viewer
accuracy
Host
accuracy
Task
accuracy
Overall
average
Ordinary multiturn SFT69.3687.7135.9968.6463.5465.7749.7051.1261.34
MA-MSFT70.0587.0636.5072.0066.4269.0949.9052.5563.01
MA-MSFT + turn-level GSPO69.9889.7928.5274.6468.8171.7650.5054.4664.42
MA-MSFT + SM-GSPO ★71.1478.8448.2079.2672.6775.5756.7158.4167.48
What changes with SM-GSPO?

Compared with MA-MSFT, trajectory-level optimization raises the overall average by 4.47 points, with clear gains in memory recall (+11.70), answer recall (+7.26), recipient routing (+6.25), and task accuracy (+5.86).

One assistant, three audiences

Useful to the entire live room

VIEWERS

Explain what matters

Real-time narration, newcomer recaps, highlights, and direct answers.

HOSTS

Surface what is emerging

Question clusters, pacing alerts, key moments, and audience feedback.

MODERATORS

Flag what needs review

Grounded risk cues, anomalies, and precise segments for follow-up.

Not talking more — knowing when to talk, what to retain, and whom to help.

Citation

Cite LiveAssistant

Read the paper on arXiv:2609.27303 ↗, and explore the official artifacts in our Hugging Face collection ↗.

@article{gao2026liveassistant,
  title   = {Live Assistant: Learning Whether, When, and Whom to
             Assist in Real-World Live Social Streams},
  author  = {Gao, Shujian and Yan, Jiamei and Yang, Yuchen and
             Zhou, Penghao and Wang, Qinglei and Fan, Tiehan and
             Wang, Yuan and Wu, Zuxuan and Jiang, Yu-Gang},
  journal = {arXiv preprint arXiv:2609.27303},
  year    = {2026},
  doi     = {10.48550/arXiv.2609.27303}
}