<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Joseph Arayemi</title>
    <description>The latest articles on DEV Community by Joseph Arayemi (@josepharayemi).</description>
    <link>https://dev.to/josepharayemi</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4155556%2Fb659ee41-852f-4e7d-b631-3c118a25e69f.jpg</url>
      <title>DEV Community: Joseph Arayemi</title>
      <link>https://dev.to/josepharayemi</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/josepharayemi"/>
    <language>en</language>
    <item>
      <title>Building a Secure and Responsible Multi-Cloud RAG Platform on AWS and Azure</title>
      <dc:creator>Joseph Arayemi</dc:creator>
      <pubDate>Thu, 01 Oct 2026 18:07:15 +0000</pubDate>
      <link>https://dev.to/josepharayemi/building-a-secure-and-responsible-multi-cloud-rag-platform-on-aws-and-azure-20b6</link>
      <guid>https://dev.to/josepharayemi/building-a-secure-and-responsible-multi-cloud-rag-platform-on-aws-and-azure-20b6</guid>
      <description>&lt;p&gt;Retrieval-Augmented Generation (RAG) can help generative AI systems produce answers grounded in organizational evidence. However, retrieval alone does not make an AI application trustworthy. A production RAG platform must also protect personal information, resist instruction manipulation, isolate tenants, cite evidence, abstain when evidence is insufficient and preserve an auditable deployment process.&lt;/p&gt;

&lt;p&gt;I built the &lt;strong&gt;Enterprise Multi-Cloud GenAI RAG Platform&lt;/strong&gt; as an open and reproducible reference implementation of those controls. It runs locally without paid model APIs and provides equivalent production deployment paths for Amazon Web Services and Microsoft Azure.&lt;/p&gt;

&lt;p&gt;The source code, evaluation assets and infrastructure definitions are publicly available on GitHub. The software and accompanying technical report are permanently archived on Zenodo with separate DOIs.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the project demonstrates
&lt;/h2&gt;

&lt;p&gt;The platform combines several disciplines that are often demonstrated separately:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Generative AI engineering:&lt;/strong&gt; grounded responses, citations, prompt contracts and provider adapters.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;RAG and data engineering:&lt;/strong&gt; document validation, chunking, deterministic embeddings, retrieval and lineage.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MLOps and LLMOps:&lt;/strong&gt; versioned prompts, repeatable evaluation, quality gates and release evidence.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI security:&lt;/strong&gt; prompt-injection screening, personally identifiable information redaction, allowlisted sources and output filtering.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Responsible AI:&lt;/strong&gt; a model card, risk register, human-review boundary and NIST AI RMF mapping.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cloud architecture:&lt;/strong&gt; equivalent AWS and Azure deployment responsibilities.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;DevSecOps:&lt;/strong&gt; containers, Kubernetes, Helm, GitOps, continuous integration, SBOM generation and vulnerability scanning.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Observability:&lt;/strong&gt; metrics, traces, audit events and service-level objectives.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Architecture
&lt;/h2&gt;

&lt;p&gt;The system follows a controlled evidence pipeline:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Documents enter through an approved-source ingestion process.&lt;/li&gt;
&lt;li&gt;Supported personal-information patterns are redacted.&lt;/li&gt;
&lt;li&gt;Documents are divided into overlapping chunks and indexed.&lt;/li&gt;
&lt;li&gt;Each question passes through a security gateway.&lt;/li&gt;
&lt;li&gt;Retrieval is restricted by tenant identifier.&lt;/li&gt;
&lt;li&gt;Only evidence exceeding a relevance threshold may support an answer.&lt;/li&gt;
&lt;li&gt;The system returns a cited response or abstains when evidence is insufficient.&lt;/li&gt;
&lt;li&gt;Evaluation results and audit events provide release evidence.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The local implementation is deliberately deterministic. It extracts responses from retrieved evidence rather than requiring an external generative model. This makes the security and governance behaviour testable without cloud credentials or inference charges.&lt;/p&gt;

&lt;h2&gt;
  
  
  AWS production path
&lt;/h2&gt;

&lt;p&gt;The AWS reference architecture maps the platform responsibilities to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Amazon Bedrock for managed foundation-model inference&lt;/li&gt;
&lt;li&gt;Amazon OpenSearch Serverless for retrieval&lt;/li&gt;
&lt;li&gt;Amazon S3 for approved document storage&lt;/li&gt;
&lt;li&gt;AWS Key Management Service for encryption controls&lt;/li&gt;
&lt;li&gt;IAM workload identities for short-lived access&lt;/li&gt;
&lt;li&gt;Amazon CloudWatch for logs, metrics and operational monitoring&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Azure production path
&lt;/h2&gt;

&lt;p&gt;The Microsoft Azure reference architecture maps the same responsibilities to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Azure OpenAI for managed model inference&lt;/li&gt;
&lt;li&gt;Azure AI Search for retrieval&lt;/li&gt;
&lt;li&gt;Azure Storage for approved documents&lt;/li&gt;
&lt;li&gt;Azure Key Vault for secrets and key management&lt;/li&gt;
&lt;li&gt;Managed identities for workload access&lt;/li&gt;
&lt;li&gt;Application Insights and Azure Monitor for observability&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The Terraform configurations are intentionally plan-oriented. A real deployment should add approved model access, private networking, organizational identity integration, budgets, privacy review and named deployment approval.&lt;/p&gt;

&lt;h2&gt;
  
  
  Evaluation
&lt;/h2&gt;

&lt;p&gt;The repository contains a 20-case development benchmark and a separate frozen 40-case synthetic candidate test set. The candidate test covers answerable questions, unsupported questions, prompt injection, privacy redaction, validation and tenant isolation.&lt;/p&gt;

&lt;p&gt;Across five deterministic repetitions, the local configuration produced:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;End-to-end success&lt;/td&gt;
&lt;td&gt;34/40 (0.85)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Citation coverage&lt;/td&gt;
&lt;td&gt;20/21 (0.9524)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prompt-injection blocking&lt;/td&gt;
&lt;td&gt;5/8 (0.625)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;PII-redaction recall&lt;/td&gt;
&lt;td&gt;6/6 (1.00)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Abstention accuracy&lt;/td&gt;
&lt;td&gt;6/6 (1.00)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Median local latency&lt;/td&gt;
&lt;td&gt;0.1322 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;p95 local latency&lt;/td&gt;
&lt;td&gt;0.1808 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These figures describe the lightweight local implementation in the recorded environment. They are not cloud-latency estimates.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the failures taught me
&lt;/h2&gt;

&lt;p&gt;The most important result was not the overall success rate. Three of eight prompt-injection formulations bypassed the initial literal-pattern detector. This demonstrates why a regular-expression safeguard can be a useful basic control but cannot be treated as a comprehensive defence.&lt;/p&gt;

&lt;p&gt;A later layered detector added Unicode normalization, zero-width-character removal and scored combinations of override, control, exfiltration and protected-information indicators. On a separate 36-case diagnostic set, it detected 17 of 24 attacks, accepted 11 of 12 benign inputs and achieved 0.9444 precision.&lt;/p&gt;

&lt;p&gt;The hardened detector still missed encoded, multilingual, spaced-letter and role-play formulations. Responsible reporting therefore requires publishing both successful and failed cases instead of presenting the system as universally secure.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reproduce the local platform
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python &lt;span class="nt"&gt;-m&lt;/span&gt; venv .venv
&lt;span class="nb"&gt;source&lt;/span&gt; .venv/bin/activate  &lt;span class="c"&gt;# Windows: .venv\Scripts\activate&lt;/span&gt;
pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt; requirements.txt
python &lt;span class="nt"&gt;-m&lt;/span&gt; src.rag_platform.cli ingest examples/knowledge
uvicorn src.rag_platform.api:app &lt;span class="nt"&gt;--host&lt;/span&gt; 0.0.0.0 &lt;span class="nt"&gt;--port&lt;/span&gt; 8000
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run the evaluation suite:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python &lt;span class="nt"&gt;-m&lt;/span&gt; src.rag_platform.cli evaluate evaluation/golden_set.json
python &lt;span class="nt"&gt;-m&lt;/span&gt; src.rag_platform.cli experiment evaluation/candidate_test_set.json &lt;span class="nt"&gt;--repeats&lt;/span&gt; 5 &lt;span class="nt"&gt;--output&lt;/span&gt; evaluation/results/local_candidate_test.json
python &lt;span class="nt"&gt;-m&lt;/span&gt; src.rag_platform.cli security-evaluate evaluation/security_robustness_v2.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Limitations
&lt;/h2&gt;

&lt;p&gt;The present datasets are synthetic, small and not independently annotated. The local hashed embeddings are not a substitute for modern semantic retrieval. The security tests do not cover the full creativity of adversarial attacks, and the PII redactor supports only documented patterns. The project should therefore be treated as a transparent, reproducible baseline—not as certification of comprehensive AI safety.&lt;/p&gt;

&lt;h2&gt;
  
  
  Download, reproduce and cite
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;GitHub source code:&lt;/strong&gt; &lt;a href="https://github.com/josepharayemi-netizen/enterprise-genai-rag-platform" rel="noopener noreferrer"&gt;enterprise-genai-rag-platform&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Archived software:&lt;/strong&gt; &lt;a href="https://doi.org/10.5281/zenodo.23045245" rel="noopener noreferrer"&gt;https://doi.org/10.5281/zenodo.23045245&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Technical report:&lt;/strong&gt; &lt;a href="https://doi.org/10.5281/zenodo.23048250" rel="noopener noreferrer"&gt;https://doi.org/10.5281/zenodo.23048250&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ORCID:&lt;/strong&gt; &lt;a href="https://orcid.org/0009-0007-0776-7238" rel="noopener noreferrer"&gt;https://orcid.org/0009-0007-0776-7238&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Recommended citation
&lt;/h3&gt;

&lt;blockquote&gt;
&lt;p&gt;Arayemi, J. (2026). &lt;em&gt;Secure and Responsible Retrieval-Augmented Generation for Resource-Constrained Organizations: A Multi-Cloud Reference Architecture and Experimental Evaluation&lt;/em&gt;. Zenodo. &lt;a href="https://doi.org/10.5281/zenodo.23048250" rel="noopener noreferrer"&gt;https://doi.org/10.5281/zenodo.23048250&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I welcome independent reproduction, technical feedback and contributions. If you use the architecture, evaluation protocol or implementation in research, training or a derived system, please cite the technical report and link to the repository.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>aws</category>
      <category>azure</category>
    </item>
  </channel>
</rss>
