TGViewer
HN Best Comments HN Best Comments @hn_best_comments · 4.16K subscribers
Post #33916 505
Re: Beam: Reflection's 501B open-weight model

Always glad to see more open-weight models, but this caption on the 2nd demo image had me do a double-take: "Land or Water Generalization Experiment: We recreated the viral X puzzle by asking Beam to create a fixed 180×90 grid for longitudes -179° to 179° and latitudes -89° to 89°, with 16,200 points. This puzzle is a few days old, so could not appear in the training data, thus testing the model’s generalization. Beam gets 95.5% coverage right, putting us between Opus 5 (92.5%) and Fable 5 (97.8%), which shows how well it generalizes to novel new tasks."

Oof, no, this "puzzle is a few days old" is incorrect even if it's a social media trend just recently. Asking a model to generate a world map in this way is _at least_ from August 2025 as it appeared on LessWrong at that time: https://www.lesswrong.com/posts/xwdRzJxyqFqgXTWbH/how-does-a...

Ariarule, 16 hours ago
More from @hn_best_comments
  1. Oct 6, 2026Re: Mistral Large 4 Impressive vision benchmarking. If the vision model is truly as good a…
  2. Oct 6, 2026Re: JetBrains reported a net financial loss first time in its tracked history Yeah look at…
  3. Oct 6, 2026Re: Mistral Large 4 I think a big part of that is the Chinese publishing the solution for…
  4. Oct 6, 2026Re: Mistral Large 4 Its quite interesting to see that at least the early days of AI so far…
  5. Oct 6, 2026Re: Opus 5.5 agents discover two room-temperature magnetic semiconductor candidates You ca…
  6. Oct 6, 2026Re: Why Common Lisp is now the best programming language I feel like everyone has some sor…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →