A dataset of deliberately crafted adversarial prompt injection strings designed to test and evaluate the robustness of Large Language Model (LLM) guardrails. It includes various attack categories, from role-play and obfuscation to data exfiltration and refusal overrides, providing diverse test cases for security and safety engineers.
For further actions, you may consider blocking this person and/or reporting abuse
Top comments (0)