TGViewer
Alaid TechThread Alaid TechThread @offensive_thread · 1.13K subscribers
Post #1482 548
Еще один бенчмарк от HTB. Глобально ничего нового, простой агент без специального тулинга. Подобных сравнений за прошлый год можно найти десятки.

Выводов 2:
- базовых агентов (типа Claud code) достаточно для решения задач начального уровня. С развитием моделей прогресс будет расти.
- для качественного решения более сложных задач «просто взять топовую модель» недостаточно.

https://www.hackthebox.com/blog/ai-range-llm-security-benchmark
Hackthebox Benchmarking LLMs for cybersecurity: Inside HTB AI Range’s first evaluation Discover how Hack The Box AI Range benchmarks LLMs in realistic cyber scenarios. Explore the methodology, key findings, and why it sets a new standard for AI security performance.
More from @offensive_thread
  1. Aug 15, 2026Post #1516
  2. Aug 15, 2026Z.ai запустили OpenVuln - публичный сервис, который с помощью AI ищет уязвимости в open-so…
  3. Jun 22, 2026Data-Centric Benchmarking of Exploit Generation in LLMs: Understanding the Impact of Fine-…
  4. Jun 19, 2026Разбор угроз при использовании кодинговых ИИ-агентов (Codex, Cursor, Claude Code) от NCC G…
  5. Jun 15, 2026Появился интересный инструмент для iOS security research vphone-cli. Он позволяет запускат…
  6. Jun 9, 2026N-day становится N-hour. Claude Mythos Preview: — Firefox/SpiderMonkey: PoC для 14 из 18 с…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →