DEV Community

Cover image for Chapter 12 — Database, APIs & Infrastructure
Black Shadow Team ©
Black Shadow Team ©

Posted on

Chapter 12 — Database, APIs & Infrastructure

#ai

Chapter 12 — Database, APIs & Infrastructure.

12.1 Introduction

The ACAI system requires a reliable infrastructure layer capable of storing users, sessions, memories, documents, embeddings, model configurations, evaluation results, and system events.

The infrastructure should be designed around three principles:

Separation of responsibilities
Security and controlled access
Scalability without unnecessary complexity

A practical prototype can begin with a small number of services and databases. As usage increases, individual components can be separated and scaled independently.

12.2 High-Level Infrastructure
USERS


┌─────────────┐
│ Frontend │
└──────┬──────┘


┌─────────────┐
│ API Gateway │
└──────┬──────┘


┌─────────────┐
│Orchestrator │
└──────┬──────┘

┌────────────────┼────────────────┐
▼ ▼ ▼
Database Retrieval Model Layer
│ │ │
▼ ▼ ▼
Metadata Vector Store LLM Server
│ │ │
└────────────────┼────────────────┘

Verification


Response
12.3 Database Architecture

ACAI should not treat every piece of information as the same kind of data.

A practical architecture separates:

Relational Data
+
Vector Data
+
Object Data
+
Cache
+
Event Data

Each storage layer has a different purpose.

12.4 Relational Database

The relational database stores structured application information.

Possible tables:

users
sessions
projects
conversations
messages
documents
models
tasks
evaluations
feedback

Example:

users

id
email
created_at
status

sessions

id
user_id
created_at
updated_at

projects

id
user_id
name
created_at

Relationships should be explicitly defined.

12.5 Conversation Storage

A conversation can be represented as:

Conversation

├── Message 1
├── Message 2
├── Message 3
└── Message 4

A message record might contain:

message_id
conversation_id
role
content
created_at
metadata

However, sensitive conversation data should not be retained indefinitely without a legitimate reason and appropriate user controls.

12.6 Memory Database

Long-term memory requires additional metadata.

Example:

memory_id
user_id
content
memory_type
importance
source
created_at
updated_at

Possible memory types:

Preference
Project Context
Important Fact
Task History
User Instruction

The system should distinguish between information explicitly provided by the user and information inferred by the model.

12.7 Vector Database

Semantic retrieval requires embeddings.

Conceptually:

Text

Embedding Model

Vector

Vector Database

Each vector should retain a reference to its source.

Vector

├── document_id
├── chunk_id
├── page
└── metadata

This allows retrieved information to be traced back to its original document.

12.8 Document Storage

Large files should generally be stored separately from the relational database.

Examples:

PDF
DOCX
Images
Audio
Video
Datasets

The database stores metadata:

document_id
filename
size
mime_type
storage_location
owner
created_at

while the actual file remains in object storage or an equivalent file-storage system.

12.9 Cache Layer

A cache can store temporary or frequently accessed information.

Possible cache entries:

Session State
Model Metadata
Rate Limits
Temporary Retrieval Results
Frequently Used Configuration

The cache should never become the only copy of information that must persist.

12.10 API Architecture

The API should expose clearly separated resources.

Example:

/api/auth
/api/chat
/api/projects
/api/documents
/api/memory
/api/search
/api/models
/api/tasks
/api/evaluation
/api/feedback
/api/health

Each API should validate requests before passing them to internal services.

12.11 Chat API

Conceptually:

POST /api/chat

Input:

{
"conversation_id": "...",
"message": "User request"
}

Processing:

Request

Authentication

Validation

Orchestrator

Response

The actual production API should also enforce rate limits and authorization.

12.12 Document API

A document pipeline could use:

POST /api/documents
GET /api/documents
GET /api/documents/{id}
DELETE /api/documents/{id}

Upload flow:

Upload

Authentication

File Validation

Malware / Safety Scan

Storage

Text Extraction

Chunking

Embedding

Indexing
12.13 Retrieval API

Example:

POST /api/search

The retrieval system can accept:

query
filters
top_k
project_id
document_scope

The service then returns ranked results.

12.14 Model API

The Model Service abstracts the underlying foundation models.

POST /internal/model/generate

The application should avoid hard-coding the rest of the system to one specific model provider.

Conceptually:

Application

Model Interface

┌──────────┬──────────┬──────────┐
│ Provider │ Provider │ Local │
│ A │ B │ Model │
└──────────┴──────────┴──────────┘

This makes model replacement easier.

12.15 Model Registry

The Model Registry stores information about available models.

Example:

model_id
provider
capabilities
context_limit
status
cost_class
version

The router can use this information when making model-selection decisions.

12.16 Task API

Complex requests may create multiple tasks.

POST /api/tasks
GET /api/tasks/{id}
POST /api/tasks/{id}/cancel

Example:

Task

├── Research
├── Analysis
├── Coding
└── Verification

Each task can have its own status:

QUEUED
RUNNING
COMPLETED
FAILED
CANCELLED
12.17 Evaluation API

The evaluation system needs its own interface.

Possible endpoints:

POST /api/evaluation/run
GET /api/evaluation/{id}
GET /api/evaluation/results

This enables automated benchmark runs without manually interacting with the application.

12.18 Authentication

The API should authenticate every protected request.

Conceptually:

User

Login

Authentication Service

Session / Token

API Request

Authorization Check

Authentication and authorization are separate:

Authentication

Who is the user?

Authorization

What is the user allowed to access?

12.19 Authorization

A user should only access resources they are permitted to access.

Example:

User A

├── Project A ✓
└── Project B ✗

Authorization should be enforced on the backend, not merely hidden in the frontend.

12.20 Rate Limiting

Without rate limiting, a single client could consume excessive resources.

Example:

User

API

Rate Limiter
├── Allowed → Continue
└── Limit → Reject / Delay

Different limits may apply to:

Authentication
Chat
File uploads
Search
Model generation
Evaluation jobs
12.21 Queue Infrastructure

Long-running tasks should not necessarily block normal API requests.

Example:

API Request


Message Queue

├── Worker 1
├── Worker 2
└── Worker 3

Workers can process:

Document ingestion
Embedding generation
Large evaluations
Video processing
Batch inference
12.22 Worker Architecture

A worker receives a job:

Job

Worker

Process

Result

Database / Event

Failed jobs can be retried according to a controlled retry policy.

12.23 GPU Infrastructure

If local or self-hosted models are used, GPU resources become important.

Conceptually:

             Model Gateway
                  │
      ┌───────────┼───────────┐
      ▼           ▼           ▼
   GPU Node 1  GPU Node 2  GPU Node 3
      │           │           │
    Model A     Model B     Model C
Enter fullscreen mode Exit fullscreen mode

The router should know which models are available on which nodes.

12.24 Horizontal Scaling

When demand increases, multiple API instances can operate simultaneously.

              Load Balancer
                   │
        ┌──────────┼──────────┐
        ▼          ▼          ▼
      API-1      API-2      API-3
Enter fullscreen mode Exit fullscreen mode

Stateless API services are easier to scale horizontally.

Persistent state should remain in shared storage systems.

12.25 Reliability

A production system should avoid single points of failure where practical.

Example:

API-1 ─┐
API-2 ─┼── Database
API-3 ─┘

If one API instance fails, the others can continue serving traffic.

The same principle can be applied to model servers and workers when resources justify it.

12.26 Backup & Recovery

Critical data should have a recovery strategy.

Primary Database


Backup


Recovery Test

A backup that has never been tested should not automatically be considered recoverable.

Important recovery metrics include:

RPO — Recovery Point Objective

How much data can potentially be lost?

RTO — Recovery Time Objective

How quickly should service be restored?

12.27 Monitoring Infrastructure

The infrastructure should monitor:

CPU
Memory
GPU
Database
API
Queue
Model
Retrieval
Latency
Errors

A centralized dashboard can provide an operational overview.

12.28 Infrastructure Security

Security should exist at multiple layers.

Internet

Firewall

Load Balancer

API

Authorization

Internal Services

Database

Sensitive credentials should be stored in secure secret-management systems rather than source code.

12.29 Development Environment

A practical development environment can initially contain:

Frontend
Backend
Database
Vector Store
Local Model / API
Testing Tools

Development and production configurations should be separated.

12.30 Production Environment

A production deployment may eventually look like:

                Internet
                   │
                   ▼
             Load Balancer
                   │
                   ▼
              API Cluster
                   │
          ┌────────┼────────┐
          ▼        ▼        ▼
       Planner   Retrieval Memory
          │        │        │
          └────────┼────────┘
                   ▼
              Model Gateway
                   │
          ┌────────┼────────┐
          ▼        ▼        ▼
        GPU-1    GPU-2    GPU-3
                   │
                   ▼
              Verification
                   │
                   ▼
                Response
Enter fullscreen mode Exit fullscreen mode

12.31 Practical Scaling Strategy

Do not begin with the full production architecture.

A realistic progression is:

Stage 1
Single Machine

Stage 2
Separate Database

Stage 3
Separate Model Server

Stage 4
Queue + Workers

Stage 5
Multiple API Instances

Stage 6
GPU Cluster

Stage 7
Multi-Region / Advanced Infrastructure

The system should only move to the next stage when actual workload requires it.

12.32 Cost Optimization

Infrastructure should be optimized according to actual measurements.

Possible strategies:

Cache repeated operations
Use smaller models for simple tasks
Batch embedding operations
Compress retrieved context
Scale workers according to demand
Shut down unused development resources

The objective is not simply minimizing cost, but finding an appropriate quality–latency–cost balance.

12.33 Infrastructure Testing

Before production, test:

API Tests
Database Tests
Retrieval Tests
Model Tests
Queue Tests
Failure Tests
Load Tests
Security Tests
Recovery Tests

A system that works with one request but fails under 100 simultaneous requests is not production-ready.

12.34 Final Infrastructure Blueprint
USERS


┌─────────────┐
│ Load/Balancer│
└──────┬──────┘


┌─────────────┐
│ API Cluster │
└──────┬──────┘


┌─────────────┐
│Orchestrator │
└──────┬──────┘

┌───────────────────┼───────────────────┐
▼ ▼ ▼
Planner Memory Retrieval
│ │ │
└───────────────────┼───────────────────┘

┌─────────────┐
│Model Router │
└──────┬──────┘

┌──────┼──────┐
▼ ▼ ▼
GPU GPU GPU
│ │ │
└──────┼──────┘

Verification


Response
12.35 Chapter Summary

The infrastructure layer provides the foundation on which the ACAI software architecture operates. A practical implementation should separate structured data, vector data, large files, cache data, and asynchronous workloads according to their specific requirements.

The most important engineering principle is progressive scaling. A prototype should begin with the smallest infrastructure capable of validating the research hypothesis. More complex infrastructure should be introduced only when measurements demonstrate a need for it.

This makes the proposed architecture technically achievable without requiring the initial project to operate as a massive AI infrastructure company.

End of Chapter 12
Stay tuned for Chapter: 13 Complete End-to-End System Architecture.

🚀 Connect with Black Shadow Team Across the Web! 🌐

We are actively sharing our latest cybersecurity research, AI safety insights, ethical hacking content, and tech updates across multiple platforms. Follow and subscribe to stay updated with our official channels:

📝 Articles & Research Papers:

Medium: https://medium.com/@blackshadowteam.net

Substack: https://blackshadowteam.substack.com

Dev.to: https://dev.to/black_shadow_team

HackerNoon: https://hackernoon.com/u/black-shadow-team

Hashnode: https://hashnode.com/@black-shadow-team

Blogspot: https://black-shadow-team.blogspot.com/

💻 Code & Open Source:

GitHub: https://github.com/blackshadowteamnet-netizen

WordPress: https://profiles.wordpress.org/blackshadowteam

📱 Social Media & Updates:

X (Twitter): https://x.com/BlackShadoTeam

Facebook Page: https://www.facebook.com/profile.php?id=61591268330812

Facebook Profile: https://www.facebook.com/profile.php?id=100090580510673

Instagram: https://www.instagram.com/black_shadow_team_x/

Threads: https://www.threads.net/@blacky_mahin_x

Bluesky: https://bsky.app/profile/black-shadow-team.bsky.social

💬 Community & Discussions:

Reddit: https://www.reddit.com/user/blackshadowteamoffic/

Quora (Bangla): https://bn.quora.com/profile/Black-Shadow-Team

Mix: https://mix.com/black_shadow_team

Discord: https://discord.com/channels/1518981404074184725/1518981404632023143

🎵 Short Videos & Audio:

TikTok: https://www.tiktok.com/@blackshadowteam.net

SoundCloud: https://on.soundcloud.com/VBWtOYsgktkw37kAza

Goodreads: https://www.goodreads.com/user/show/203582586-black-shadow-team-team

Stay connected and join our growing cybersecurity community! 🛡️✨

Top comments (0)