TGViewer
DevOps MemOps DevOps MemOps @devops_memops · 6.34K subscribers
Post #7233 1.64K
Autoscaling Hid Our LLM Cost Regression (85% → 4% Cache Hit Rate)

Почему мы перенесли проектирование мощностей в CI и начали ограничивать эффективность кэширования префиксов

📌 Подробнее: https://medium.com/@nroan/autoscaling-hid-our-llm-cost-regression-85-4-cache-hit-rate-b4beab5df240

MemOps 🤨
Medium Autoscaling Hid Our LLM Cost Regression (85% → 4% Cache Hit Rate) Why we moved capacity engineering into CI and started gating on prefix-cache efficiency
More from @devops_memops
  1. Oct 4, 2026MemOps 😃
  2. Oct 4, 2026Post #8344
  3. Oct 3, 2026whiteboard — приложение для пк с открытым исходным кодом, в котором люди и агенты могут со…
  4. Oct 3, 2026MemOps 😃
  5. Oct 3, 2026ClickHouse под трейсы: почему шесть бэкендов Jaeger оказались выбором из двух В документац…
  6. Oct 2, 2026MemOps 😃
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →