Your AI agent can execute code, access databases, and call APIs.
But can it recognize malicious instructions?
Prompt injection. Memory poisoning. Tool misuse. Secret leaks.
Are you testing for any of these?
I wrote a beginner-to-pro guide:
Your AI agent can execute code, access databases, and call APIs.
But can it recognize malicious instructions?
Prompt injection. Memory poisoning. Tool misuse. Secret leaks.
Are you testing for any of these?
I wrote a beginner-to-pro guide:
For further actions, you may consider blocking this person and/or reporting abuse
Top comments (1)
Bài viết chạm đúng vào nỗi đau thực tế khi triển khai agent vào production. Mình thấy 3 lỗ hổng thường bị bỏ qua:
Prompt injection qua tool output — Agent gọi API trả về data chứa instruction ẩn, model tuân theo mà không kiểm tra. Cần sanitize output của mọi tool trước khi feed lại vào context.
Privilege escalation qua function calling — Cho agent quyền
execute_sqlhaycall_apimà không có policy-based access control ở layer tool. Nên wrap mỗi tool bằng adapter kiểm tra scope/role trước khi execute.Data exfiltration qua chained calls — Agent bị dẫn dắt gọi tool A lấy token, tool B dùng token đó truy cập resource nhạy cảm. Cần audit trail đầy đủ + rate limit per session, không chỉ per API key.
Mình đang dùng approach: tách hoàn toàn planning (LLM) và execution (deterministic sandbox). Planner chỉ output plan JSON, executor validate từng step against policy engine mới chạy. Giảm surface attack đáng kể.
Curious: team bạn handle việc versioning prompt/policy như thế nào khi rollback agent behavior? PS: the tool I meant is on labagent .tech