DEV Community

Cover image for Building AI Agents? Here's What Can Go Wrong.
Arunima Chaudhuri
Arunima Chaudhuri

Posted on

Building AI Agents? Here's What Can Go Wrong.

Your AI agent can execute code, access databases, and call APIs.

But can it recognize malicious instructions?

Prompt injection. Memory poisoning. Tool misuse. Secret leaks.

Are you testing for any of these?

I wrote a beginner-to-pro guide:

AgentSec 101 →

Top comments (1)

Collapse
 
koev3kcjausd profile image
koev3kcjausd •

Bài viết chạm đúng vào nỗi đau thực tế khi triển khai agent vào production. Mình thấy 3 lỗ hổng thường bị bỏ qua:

  1. Prompt injection qua tool output — Agent gọi API trả về data chứa instruction ẩn, model tuân theo mà không kiểm tra. Cần sanitize output của mọi tool trước khi feed lại vào context.

  2. Privilege escalation qua function calling — Cho agent quyền execute_sql hay call_api mà không có policy-based access control ở layer tool. Nên wrap mỗi tool bằng adapter kiểm tra scope/role trước khi execute.

  3. Data exfiltration qua chained calls — Agent bị dẫn dắt gọi tool A lấy token, tool B dùng token đó truy cập resource nhạy cảm. Cần audit trail đầy đủ + rate limit per session, không chỉ per API key.

Mình đang dùng approach: tách hoàn toàn planning (LLM) và execution (deterministic sandbox). Planner chỉ output plan JSON, executor validate từng step against policy engine mới chạy. Giảm surface attack đáng kể.

Curious: team bạn handle việc versioning prompt/policy như thế nào khi rollback agent behavior? PS: the tool I meant is on labagent .tech