TGViewer
Parallel Experiments Parallel Experiments @linghaoch · 1.75K subscribers
Post #1012 1.14K
https://alexzhang13.github.io/blog/2026/harness/

> A good harness is a harness that reduces unfamiliar problems to familiar ones and reduces complex problems to simple ones. In other words, even if the state s is out-of-distribution (OOD) to what any individual language model call was trained for, a good harness produces observations o that are locally in-distribution (LID), which we define as every individual LM call over this observation being in-distribution with respect to the training data.
Alex L. Zhang Language model harnesses are compositional generalizers Harnesses can lead to compositional generalization: we observe a property in training RLMs, in which similarly structured tasks are viewed as isomorphic and all individual LM calls in the harness become in-distribution.
  • 👀 1
More from @linghaoch
  1. Oct 1, 2026https://56k.rip/
  2. Sep 11, 2026https://sunilpai.dev/posts/the-task-isnt-the-job/
  3. Jul 6, 2026https://linghao.io/posts/taxonomy-differences-matter 以前觉得 taxonomy 只是无聊的分类学,开始做 LLM qualit…
  4. Jun 28, 2026关于层出不穷的各式 AI memory system 的一些思考:我们应该把更多的精力放在设计更好的 eval 上,从而让最强的 memory system 进化出来 https:…
  5. May 17, 2026https://arxiv.org/abs/2503.02113 The core idea: Deep learning does not work because neural…
  6. Apr 18, 2026一月底最后一个周六有了一个灵感,想做个解放双手,优化了 AirPods 录音,边散步边和自己对话的 App。打开 Cursor coding 了一天,第二天就出门去 SoHo 散步…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →