<obs>Observe
Keep watching without disturbing the stream when no intervention is needed.
Selective proactive assistance as a decision trajectory
Live Assistant learns whether to act, when assistance is needed, and whom to address — turning passive video understanding into timely participation.
Instead of waiting for a prompt, LiveAssistant continuously reads the social stream and chooses whether to observe, remember, or assist the right participant.
Each frame is one step in a continuous trajectory. The model can wait, preserve context, or produce a role-aware response only when the stream calls for it.
This complete 10-minute example visualizes all 60 causal decisions, from state selection to recipient routing, and explains why useful assistance sometimes means not emitting anything at all.
Most streaming models are optimized to keep talking. LiveAssistant reframes livestream understanding as selective participation in a shared, multi-party environment.
<obs>Keep watching without disturbing the stream when no intervention is needed.
<mem>Write a private clue to memory so future decisions retain causal context.
<ans>Speak to viewers, hosts, or moderators with a specific, grounded task.
A native livestream task requires more than frame understanding. It must reason causally over time, route actions across roles, and learn restraint.
Causal, mixed-initiative decision-making over heterogeneous, long-horizon signals.
Viewer, host, and moderator are first-class policy decisions.
Video, audio, comments, gifts, dynamics, and metadata are fused directly.
Aligned 10-second decisions with automatic checks and human review.
MA-MSFT learns the grammar; SM-GSPO optimizes the streaming policy.
The same decision trajectory connects annotation, supervised fine-tuning, reinforcement learning, and bounded-cache inference.

Align video, audio, interaction, gifts, dynamics, and metadata in causal order.
MA-MSFT protects rare structural decisions with marker-aware objectives.
SM-GSPO combines structure, content, turn-level, and trajectory-level rewards.
A dense recent window and compressed long-term memory preserve useful context.
Evaluation uses native audio, video, comments, and gifts. The non-trivial state distribution penalizes both “always talk” and “always silent” policies.
| Training strategy | State overall | OBS recall | MEM recall | ANS recall | Recipient overall | Viewer accuracy | Host accuracy | Task accuracy | Overall average |
|---|---|---|---|---|---|---|---|---|---|
| Ordinary multiturn SFT | 69.36 | 87.71 | 35.99 | 68.64 | 63.54 | 65.77 | 49.70 | 51.12 | 61.34 |
| MA-MSFT | 70.05 | 87.06 | 36.50 | 72.00 | 66.42 | 69.09 | 49.90 | 52.55 | 63.01 |
| MA-MSFT + turn-level GSPO | 69.98 | 89.79 | 28.52 | 74.64 | 68.81 | 71.76 | 50.50 | 54.46 | 64.42 |
| MA-MSFT + SM-GSPO ★ | 71.14 | 78.84 | 48.20 | 79.26 | 72.67 | 75.57 | 56.71 | 58.41 | 67.48 |
Compared with MA-MSFT, trajectory-level optimization raises the overall average by 4.47 points, with clear gains in memory recall (+11.70), answer recall (+7.26), recipient routing (+6.25), and task accuracy (+5.86).
Real-time narration, newcomer recaps, highlights, and direct answers.
Question clusters, pacing alerts, key moments, and audience feedback.
Grounded risk cues, anomalies, and precise segments for follow-up.
Not talking more — knowing when to talk, what to retain, and whom to help.
Read the paper on arXiv:2609.27303 ↗, and explore the official artifacts in our Hugging Face collection ↗.
@article{gao2026liveassistant,
title = {Live Assistant: Learning Whether, When, and Whom to
Assist in Real-World Live Social Streams},
author = {Gao, Shujian and Yan, Jiamei and Yang, Yuchen and
Zhou, Penghao and Wang, Qinglei and Fan, Tiehan and
Wang, Yuan and Wu, Zuxuan and Jiang, Yu-Gang},
journal = {arXiv preprint arXiv:2609.27303},
year = {2026},
doi = {10.48550/arXiv.2609.27303}
}