TGViewer
All about AI, Web 3.0, BCI All about AI, Web 3.0, BCI @alwebbci · 3.9K subscribers
Post #4380 632
Great paper from Google.
Why context beats scale for agents working against unfamiliar APIs.

GPU kernel optimization has KernelBench to hillclimb on. TPUs had nothing, and the Pallas DSL is documented thinly enough that models mostly guess.

New research from Google, Harvard, and UC Berkeley introduces JAXBench, 50 JAX workloads built from real MaxText architectures like Llama-3.1, DeepSeek-V3, Mixtral, Mamba-2, and AlphaFold2.

Eight operators ship with hand-tuned Pallas kernels from Tokamax, so agent output gets measured against expert work instead of a naive baseline.

With Gemini 3 Flash, conditioning on curated TPU documentation lifts per-sample correctness from 5.8% to 37.3%, solving 48 of 50 benchmarks at a 1.28x geomean speedup. Beam search then pushes it to 1.36x.

Correctness turned out to be a documentation problem and speed turned out to be a search problem.
arXiv.org JAXBench: Benchmarking Autonomous TPU Kernel Optimization Rigorous benchmarks have driven progress in autonomous GPU kernel performance optimization by establishing a shared target to hillclimb on, but no equivalent exists for TPUs. We present JAXBench,...
  • ❤ 6
  • 🔥 2
  • 🥰 2
More from @alwebbci
  1. Oct 2, 2026Google research introduced a next-generation Federated Learning system. By using Trusted E…
  2. Oct 2, 2026Apple introduced LoopCD Heard the rumors that frontier models like GPT-6 Astra and Gemini…
  3. Oct 2, 2026Anthropic just published its 𝗳𝘂𝗹𝗹 𝗴𝗲𝘁𝘁𝗶𝗻𝗴-𝘀𝘁𝗮𝗿𝘁𝗲𝗱 𝗴𝘂𝗶𝗱𝗲 for Claude…
  4. Oct 1, 2026DeepSeek released Harness v0.2 preview, with a desktop app available for macOS and Windows…
  5. Oct 1, 2026Meet Gemini 4 Argon It’s built for complex workflows across coding, enterprise knowledge w…
  6. Sep 30, 2026Meta presents Context Language Models Current harnesses have a very rigid setup in how the…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →