🔬 ML PAPERS (arXiv, Feb 19)
📄 Towards a Science of AI Agent Reliability — 12 metrics for agent eval beyond success rate
→ arxiv.org/abs/2602.16666
📄 Agent Skill Framework — Can SLMs benefit from agent skills (Copilot, LangChain style)? On-prem focus
→ arxiv.org/abs/2602.16653
📄 Framework of Thoughts (FoT) — Unifies CoT/ToT/GoT with dynamic reasoning + auto-tuning
→ arxiv.org/abs/2602.16512
📄 Calibrate-Then-Act — LLM agents that reason about cost vs uncertainty tradeoffs
→ arxiv.org/abs/2602.16699
📄 MMA: Multimodal Memory Agent — Reliability scoring for retrieved memories (decay, credibility, conflict)
→ arxiv.org/abs/2602.16493
📄 Self-Supervised Semantic Bridge — Diffusion bridge for unpaired img2img translation + text-guided editing
→ arxiv.org/abs/2602.16664
📄 TeCoNeRV — Neural video compression with 20× memory reduction, +5.35dB PSNR at 720p
→ arxiv.org/abs/2602.16711
📄 ReMoRa — Long-video MLLM using compressed representations (sparse keyframes + motion)
→ arxiv.org/abs/2602.16412
Post #4914
3.11K