Machine learningLLMs & Text
SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding
The paper introduces SafeDecoding, a strategy to protect large language models from jailbreak attacks, maintaining the quality of responses to user queries while reducing attack success rate.
Featured in No. 53 on 12 Jun 2024 · · 300 citations today · published in Annual Meeting of the Association for Computational Linguistics
- Released
- 14 Feb 2024
- First featured
- No. 53 · 12 Jun 2024
- Citations (Semantic Scholar)
- 300
- Influential citations
- 45
- Published in
- Annual Meeting of the Association for Computational Linguistics
- Shares when featured
- 103
- Identifier
- arXiv:2402.08983
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).