TGViewer
ML&|Sec Feed ML&|Sec Feed @mlsecfeed · 1.38K subscribers
Post #2023 188

Forwarded from CyberSecurityTechnologies

NIST_AI_800-2.pdf681.5 KB
#MLSecOps
#Infosec_Standards
NIST AI 800-2 (IPD):
"Practices for Automated Benchmark Evaluations of Language Models", Jan. 2026.

// This document identifies practices for conducting automated benchmark evaluations of language models and similar general-purpose AI models that output text. Evaluations of these models, often embedded into systems capable of functioning as chatbots and AI agents, are increasingly common. However, consistent practices to support the validity and reproducibility of such evaluations are only beginning to emerge. The practices presented in this document are intended to reflect best practices; where relevant, practices that are relatively less mature in ecosystem use are labeled as emerging practice
  • ✍ 2
More from @mlsecfeed
  1. Oct 6, 2026https://huggingface.co/spaces/AlexWortega/openjev чат а го накидаем лайков жеска, я чет за…
  2. Oct 6, 2026Bitdefender внезапно выпустила свой бесплатный антивирус для ИИ-агентов 💪 AI Guardian сле…
  3. Oct 5, 2026https://poloclub.github.io/transformer-explainer/
  4. Oct 4, 2026GLM-5.3 без цензуры: вышел EXL3-квант на 3 бита На Hugging Face появилась uncensored-верси…
  5. Oct 3, 2026На Хабре вышла статья о том, что ИИ изменил в кибератаках от нашей Red Team команды (tl;dr…
  6. Oct 2, 2026Cloudflare тоже сделали свои версии Jev Они выпустили Clef и Clef-flash, две открытые deci…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →