TGViewer
Github LLMs Github LLMs @llm_learning · 757 subscribers
Post #62 1.44K
Recent explorations with commercial Large Language Models (LLMs) have shown that non-expert users can jailbreak LLMs by simply manipulating their prompts; resulting in degenerate output behavior, privacy and security breaches, offensive outputs, and violations of content regulator policies. Limited studies have been conducted to formalize and analyze these attacks and their mitigations. We bridge this gap by proposing a formalism and a taxonomy of known (and possible) jailbreaks. We survey existing jailbreak methods and their effectiveness on open-source and commercial LLMs (such as GPT-based models, OPT, BLOOM, and FLAN-T5-XXL). We further discuss the challenges of jailbreak detection in terms of their effectiveness against known attacks. For further analysis, we release a dataset of model outputs across 3700 jailbreak prompts over 4 tasks.

🗂 Paper: https://arxiv.org/pdf/2305.14965

@scopeofai
@LLM_learning
  • 🔥 3
More from @llm_learning
  1. Sep 11, 2026document post
  2. Sep 11, 2026Large-language models (LLMs) are rapidly becoming part of human culture, reshaping how inf…
  3. Jun 29, 2026CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation https://arx…
  4. Oct 2, 2025When AIs Judge AIs: The Rise of Agent-as-a-Judge Evaluation for LLMs 📚 Read @LLM_learning
  5. Aug 5, 2025With so many LLM papers being published, it's hard to keep up and compare results. This st…
  6. Jun 16, 2025Deep-Live-Cam Real time face swap and one-click video deepfake with only a single image Cr…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →