论文

MOBA-VL:实时MOBA赛事解说模型

MOBA-VL: Event-Localized Multi-Turn Reinforcement Learning for Real-Time MOBA Commentary

精选理由

MOBA-VL用游戏遥测数据训练,能准确捕捉击杀和目标事件,比现有解说模型更精准。

MOBA-VL是一个90亿参数的视觉语言模型,专为实时MOBA赛事解说设计。该模型使用游戏遥测数据作为监督信号,通过事件定位多轮强化学习训练。在MOBACast-Bench基准测试中,MOBA-VL在完整赛事解说上达到63.25分,高于StreamingVLM的55.12分;在片段解说上达到63.45分,高于DeepSeek-V4.1-Flash的56.22分。事件定位训练使事件召回率从34.5提升至42.1。

原文 · arXiv: DeepSeek

MOBA-VL: Event-Localized Multi-Turn Reinforcement Learning for Real-Time MOBA Commentary

Real-time commentary for Multiplayer Online Battle Arena (MOBA) esports requires a vision-language model (VLM) to narrate a live match second by second, both fluently and accurately. Existing streaming VLMs sound natural but often miss key events such as kills and objectives. To address this limitation, we use game telemetry, which records exactly when each event occurs, as a supervision signal. We introduce MOBA-VL, a 9B-parameter model trained on this signal with event-localized multi-turn reinforcement learning, which rewards the turns that describe each event. We also collect MOBACast, 860 professional matches (about 460 hours) across three MOBA games with word-level timestamped commentary, and MOBACast-Bench, a benchmark from held-out tournaments. On MOBACast-Bench, MOBA-VL achieves the highest Overall score on full matches (63.25 vs. 55.12 for StreamingVLM) and clips (63.45 vs. 56.22 for DeepSeek-V4.1-Flash). Event-localized credit also raises event recall from 34.5 to 42.1 over supervised fine-tuning. Code and data will be released, and demos are available on an anonymous project page at https://moba-vl.github.io.