TGViewer
Parallel Experiments Parallel Experiments @linghaoch · 1.75K subscribers
Post #998 835
https://si.inc/posts/fdm1/

这个新的 computer use model 有点厉害,号称解决了两个难点:

1. 高质量的有监督视频数据是稀缺的,scale 上不去。

解决方案:先用少量有监督数据训练一个 inverse dynamics model(根据视频帧数据预测键鼠输入是什么),再用它去标注了 1100 万个小时的视频数据。

2. video encoder 效率不高,vlm 经常耗费大量 token 只能处理几秒钟的 30 fps 视频输入。

解决方案:注意到为 computer use model 所做的视频标注本就是 non causal 的(你得看到视频上打出字来才能知道键盘按了什么),于是基于 masked diffusion 架构去训练 video encoder,最终效率达到了惊人的 1 million token 可以编码 2 小时 30 fps 的视频。

解决这两点使得最终模型的训练得以 scale 到一个前所未有的程度。
blog The First Fully General Computer Action Model We trained a model on our 11-million-hour video dataset. Our model can explore complex websites, complete multi-action CAD modeling sequences, and drive a car in the real world, all at 30 FPS.
More from @linghaoch
  1. Oct 1, 2026https://56k.rip/
  2. Sep 11, 2026https://sunilpai.dev/posts/the-task-isnt-the-job/
  3. Jul 25, 2026https://alexzhang13.github.io/blog/2026/harness/ > A good harness is a harness that reduce…
  4. Jul 6, 2026https://linghao.io/posts/taxonomy-differences-matter 以前觉得 taxonomy 只是无聊的分类学,开始做 LLM qualit…
  5. Jun 28, 2026关于层出不穷的各式 AI memory system 的一些思考:我们应该把更多的精力放在设计更好的 eval 上,从而让最强的 memory system 进化出来 https:…
  6. May 17, 2026https://arxiv.org/abs/2503.02113 The core idea: Deep learning does not work because neural…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →