Stability AI 发布 SemanTok:让视频世界模型更小但更高效
Stability AI 出了个叫 SemanTok 的新方法,把视频开头的 token 做得更语义化,小模型能打三倍大的模型,做视频生成的可以看看这篇论文。
Stability AI 的 Interactive Research 团队发布 SemanTok 方法,针对视频生成中从粗到细的 token 排列方式,让前几个 token 携带更多语义信息,使模型对场景的表征更容易预测。论文指出,这样做能让视频生成的效率更高。实验结果显示,使用 SemanTok 的模型在性能上追平或超过参数量三倍以上的模型。
What if making video-based world models smaller isn't just about better compression, but about making their representations easier to predict?
Recent approaches to video generation explore building scenes from coarse to fine. The first few tokens capture the big picture, like a person playing a guitar, while later tokens progressively add visual details.
Our Interactive Research team just published SemanTok, which takes this idea further. By making those early tokens more semantically meaningful, we give the model a clearer understanding of what's happening in a scene, making the representation easier to predict and video generation more efficient.
The result: a model using SemanTok matches or beats the performance of a model more than three times its size.
Read the full paper: https://t.co/uBHxvEaAi9