<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: CourtGPT</title>
    <description>The latest articles on DEV Community by CourtGPT (@courtgpt).</description>
    <link>https://dev.to/courtgpt</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4047437%2Fe94a093b-43aa-4d57-a1ec-a82806484f91.jpg</url>
      <title>DEV Community: CourtGPT</title>
      <link>https://dev.to/courtgpt</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/courtgpt"/>
    <language>en</language>
    <item>
      <title>OpenAI Cookbook RAG Best Practices: A Study Summary</title>
      <dc:creator>CourtGPT</dc:creator>
      <pubDate>Sun, 26 Jul 2026 02:18:59 +0000</pubDate>
      <link>https://dev.to/courtgpt/10-ai-engineering-portfolio-projects-that-land-200k-jobs-42no</link>
      <guid>https://dev.to/courtgpt/10-ai-engineering-portfolio-projects-that-land-200k-jobs-42no</guid>
      <description></description>
    </item>
    <item>
      <title>State Bar AI Ethics Opinions: A 2026 Reference Guide</title>
      <dc:creator>CourtGPT</dc:creator>
      <pubDate>Sun, 26 Jul 2026 02:16:32 +0000</pubDate>
      <link>https://dev.to/courtgpt/why-i-recommend-webflow-shopify-for-ai-pohow-i-built-courtgptai-a-2m-legal-ai-agent-serving-14kg</link>
      <guid>https://dev.to/courtgpt/why-i-recommend-webflow-shopify-for-ai-pohow-i-built-courtgptai-a-2m-legal-ai-agent-serving-14kg</guid>
      <description>&lt;h1&gt;
  
  
  State Bar AI Ethics Opinions: A 2026 Reference Guide
&lt;/h1&gt;

&lt;p&gt;Lawyers increasingly rely on AI for document review, research, and drafting. State bars have responded at different speeds. This guide summarizes published AI ethics opinions as of July 2026.&lt;/p&gt;

&lt;h2&gt;
  
  
  ABA Standing Committee on Ethics
&lt;/h2&gt;

&lt;p&gt;Formal Opinion 512 (July 2024): Lawyers must supervise non-lawyer assistance, including AI. Substantially equivalent to supervising junior associates.&lt;/p&gt;

&lt;p&gt;Source: &lt;a href="https://www.americanbar.org/content/dam/aba/publications/dispbullet/2024-july-2024/512.pdf" rel="noopener noreferrer"&gt;https://www.americanbar.org/content/dam/aba/publications/dispbullet/2024-july-2024/512.pdf&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Formal Opinion 533 (December 2024): A lawyers duty of technology competence extends to understanding generative AIs benefits AND risks (hallucination, bias, confidentiality). Lawyers must advise clients about AIs role in legal work.&lt;/p&gt;

&lt;p&gt;Source: &lt;a href="https://www.americanbar.org/content/dam/aba/publications/dispbullet/2024-december-2024/533.pdf" rel="noopener noreferrer"&gt;https://www.americanbar.org/content/dam/aba/publications/dispbullet/2024-december-2024/533.pdf&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  State-Level Opinions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  California (2024)
&lt;/h3&gt;

&lt;p&gt;A lawyer may use AI in the practice of law, provided the lawyer uses AI as a tool with the lawyers informed analysis and independent judgment.&lt;/p&gt;

&lt;p&gt;Source: &lt;a href="https://www.calbar.ca.gov/Portals/0/documents/Final-Practical-Guidance-for-the-Use-of-Generative-Artificial-Intelligence-in-the-Practice-of-Law.pdf" rel="noopener noreferrer"&gt;https://www.calbar.ca.gov/Portals/0/documents/Final-Practical-Guidance-for-the-Use-of-Generative-Artificial-Intelligence-in-the-Practice-of-Law.pdf&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Florida (2024)
&lt;/h3&gt;

&lt;p&gt;Lawyers must (1) maintain competence in AI technology use, (2) protect client confidentiality, (3) avoid deceptive practices, (4) bill reasonably for AI-assisted tasks.&lt;/p&gt;

&lt;p&gt;Source: &lt;a href="https://www.floridabar.org/the-florida-bar-news/florida-bar-board-of-governors-issues-advisory-opinion-on-generative-ai/" rel="noopener noreferrer"&gt;https://www.floridabar.org/the-florida-bar-news/florida-bar-board-of-governors-issues-advisory-opinion-on-generative-ai/&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  New York (2025)
&lt;/h3&gt;

&lt;p&gt;The use of generative AI is consistent with a lawyers obligation to provide competent representation.&lt;/p&gt;

&lt;p&gt;NYSBA Opinion 24-87: &lt;a href="https://nysba.org/wp-content/uploads/2025/01/Opinion-24-87-GAI-Use.pdf" rel="noopener noreferrer"&gt;https://nysba.org/wp-content/uploads/2025/01/Opinion-24-87-GAI-Use.pdf&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  New Jersey (2025)
&lt;/h3&gt;

&lt;p&gt;Lawyers may use AI but must (1) understand limitations, (2) supervise AI work, (3) verify citations, (4) protect confidentiality, (5) communicate AI use to clients.&lt;/p&gt;

&lt;p&gt;Source: &lt;a href="https://www.njcourts.gov/courts/assets/courts/supreme/aos/advance-sheet-opinions/2025/njs17823.pdf" rel="noopener noreferrer"&gt;https://www.njcourts.gov/courts/assets/courts/supreme/aos/advance-sheet-opinions/2025/njs17823.pdf&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Texas (2025)
&lt;/h3&gt;

&lt;p&gt;Lawyers using AI in litigation must disclose the technology used.&lt;/p&gt;

&lt;p&gt;Source: &lt;a href="https://www.legalethicstexas.com/wp-content/uploads/2025/07/2025-Adv-Op-05.pdf" rel="noopener noreferrer"&gt;https://www.legalethicstexas.com/wp-content/uploads/2025/07/2025-Adv-Op-05.pdf&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Common themes
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Competence required (ABA Model Rule 1.1)&lt;/li&gt;
&lt;li&gt;Verification of citations mandatory&lt;/li&gt;
&lt;li&gt;Confidentiality with BAA-protected tools&lt;/li&gt;
&lt;li&gt;Reasonable fees for AI-assisted work&lt;/li&gt;
&lt;li&gt;Communication with clients&lt;/li&gt;
&lt;li&gt;Disclosure to court (FRCP Rule 26 in federal courts)&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;ABA Formal Opinion 512 (2024)&lt;/li&gt;
&lt;li&gt;ABA Formal Opinion 533 (2024)&lt;/li&gt;
&lt;li&gt;California State Bar (2024)&lt;/li&gt;
&lt;li&gt;Florida Bar (2024)&lt;/li&gt;
&lt;li&gt;NYSBA Opinion 24-87 (2025)&lt;/li&gt;
&lt;li&gt;New Jersey AI Opinion (2025)&lt;/li&gt;
&lt;li&gt;Texas AI Opinion (2025)&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  About
&lt;/h2&gt;

&lt;p&gt;Dillon Deutsch built CourtGPT.ai for AI-augmented legal research with citation verification. &lt;a href="https://courtgpt.ai" rel="noopener noreferrer"&gt;https://courtgpt.ai&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This article is informational and does not constitute legal advice. Consult your state bars full opinion text for binding guidance.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>n8n Workflow Automation: A Reference Architecture for Legal and Business Operations</title>
      <dc:creator>CourtGPT</dc:creator>
      <pubDate>Sun, 26 Jul 2026 01:50:47 +0000</pubDate>
      <link>https://dev.to/courtgpt/5-n8n-automation-patterns-that-saved-my-clients-20-hours-per-week-5ha3</link>
      <guid>https://dev.to/courtgpt/5-n8n-automation-patterns-that-saved-my-clients-20-hours-per-week-5ha3</guid>
      <description>&lt;h1&gt;
  
  
  n8n Workflow Automation: A Reference Architecture for Legal and Business Operations
&lt;/h1&gt;

&lt;p&gt;n8n is an open-source workflow automation tool used by 100,000+ organizations as of 2026. This article summarizes documented n8n patterns useful for legal operations, with attention to data handling, audit trails, and reliability constraints that matter in regulated contexts.&lt;/p&gt;

&lt;h2&gt;
  
  
  What n8n is
&lt;/h2&gt;

&lt;p&gt;n8n is a node-based workflow builder that connects APIs, databases, and SaaS tools. Key properties:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Open-source core (Apache 2.0 from v0.224+)&lt;/li&gt;
&lt;li&gt;400+ nodes for common integrations&lt;/li&gt;
&lt;li&gt;Self-hosted or hosted (n8n.cloud)&lt;/li&gt;
&lt;li&gt;JavaScript and Python execution in code nodes&lt;/li&gt;
&lt;li&gt;Webhook, polling, and scheduled triggers&lt;/li&gt;
&lt;li&gt;Queue mode for production scale&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Source: &lt;a href="https://docs.n8n.io/" rel="noopener noreferrer"&gt;https://docs.n8n.io/&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Patterns
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Webhook + Filter + Action (3-node pipeline)
&lt;/h3&gt;

&lt;p&gt;Trigger: incoming webhook (e.g., form submission, CRM event).&lt;br&gt;
Filter: optional condition check (e.g., "value &amp;gt; threshold").&lt;br&gt;
Action: API call to downstream system.&lt;/p&gt;

&lt;p&gt;Use cases:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Form submission -&amp;gt; CRM add&lt;/li&gt;
&lt;li&gt;Stripe payment event -&amp;gt; Slack notification&lt;/li&gt;
&lt;li&gt;GitHub issue opened -&amp;gt; Linear task&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Reliability notes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Webhook endpoints must respond 200 within 5s (use "Respond to Webhook" node)&lt;/li&gt;
&lt;li&gt;Idempotency keys required for retries&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  2. Database sync (Postgres -&amp;gt; Postgres)
&lt;/h3&gt;

&lt;p&gt;Use a Postgres trigger node reading from one DB and writing to another, with column mapping.&lt;/p&gt;

&lt;p&gt;Use cases:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Replicate CRM data to analytics warehouse&lt;/li&gt;
&lt;li&gt;Sync legal case events to compliance audit DB&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Reliability notes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Use &lt;code&gt;UPSERT&lt;/code&gt; semantics to handle race conditions&lt;/li&gt;
&lt;li&gt;Track last-sync timestamp in a control table&lt;/li&gt;
&lt;li&gt;Handle null values explicitly in mapping&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  3. RAG retrieval pipeline
&lt;/h3&gt;

&lt;p&gt;Trigger: chat input or scheduled poll.&lt;br&gt;
Action 1: Query vector DB (Pinecone, pgvector, Weaviate).&lt;br&gt;
Action 2: Call OpenAI / Anthropic with retrieved context.&lt;br&gt;
Action 3: Send response to chat platform (Slack, MS Teams, Discord).&lt;/p&gt;

&lt;p&gt;Reliability notes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Set timeouts on each node (recommended: 30s for LLM call)&lt;/li&gt;
&lt;li&gt;Implement fallback to "I don't know" if retrieval fails&lt;/li&gt;
&lt;li&gt;Log every query for compliance audits&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  4. Document processing chain
&lt;/h3&gt;

&lt;p&gt;Trigger: document upload to S3.&lt;br&gt;
Action 1: Textract/equivalent OCR.&lt;br&gt;
Action 2: Classify with LLM.&lt;br&gt;
Action 3: Route to approval queue (Slack, Notion, email).&lt;/p&gt;

&lt;p&gt;Use cases:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Contract intake and classification&lt;/li&gt;
&lt;li&gt;Invoice processing&lt;/li&gt;
&lt;li&gt;Patient onboarding forms (with BAA-protected LLM)&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  5. Multi-step approval workflow
&lt;/h3&gt;

&lt;p&gt;Trigger: form submission.&lt;br&gt;
Action 1: Auto-validate inputs.&lt;br&gt;
Action 2: Send to approver (Slack).&lt;br&gt;
Action 3: On approval, log + execute downstream action.&lt;br&gt;
Action 4: Notify requester.&lt;/p&gt;

&lt;p&gt;Reliability notes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;All decisions must be logged for audit&lt;/li&gt;
&lt;li&gt;Include timestamp + actor + decision reason in log&lt;/li&gt;
&lt;li&gt;Implement timeout (e.g., 7 days, then escalate)&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Production hardening checklist
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;[ ] Use queue mode for multi-instance scale&lt;/li&gt;
&lt;li&gt;[ ] Configure credentials as separate &lt;code&gt;.env&lt;/code&gt; files, not hardcoded&lt;/li&gt;
&lt;li&gt;[ ] Set per-node timeouts (recommended: 30-60s)&lt;/li&gt;
&lt;li&gt;[ ] Implement retry with exponential backoff&lt;/li&gt;
&lt;li&gt;[ ] Add error workflow to send failures to ops channel&lt;/li&gt;
&lt;li&gt;[ ] Enable workflow-level logging&lt;/li&gt;
&lt;li&gt;[ ] Use webhook IP allowlists where applicable&lt;/li&gt;
&lt;li&gt;[ ] Comply with client data residency (US/EU-only deployments)&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Compliance considerations
&lt;/h2&gt;

&lt;p&gt;For law firms, legal aid clinics, and in-house legal teams using n8n with client data:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;HIPAA&lt;/strong&gt;: BAA-protected LLM endpoints only (e.g., Azure OpenAI, AWS Bedrock with BAA)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GDPR&lt;/strong&gt;: EU data residency for processing EU client data; document lawful basis&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Privilege considerations&lt;/strong&gt;: Avoid sending privileged communications through n8n webhooks to non-privileged systems without bar review&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Audit logging&lt;/strong&gt;: All workflow executions must be retained for 7+ years&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Data minimization&lt;/strong&gt;: Workflows should not pull more client data than needed&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Limitations and gotchas
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Hard limits&lt;/strong&gt;: 25 nodes per workflow (community edition); 1000 runs in queue mode retention&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;State handling&lt;/strong&gt;: Built-in state is per-workflow, not globally; use external DB for cross-workflow state&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Concurrency&lt;/strong&gt;: Free tier limits concurrency; queue mode handles this&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Custom code&lt;/strong&gt;: Code nodes have execution timeouts; long-running operations should be broken up&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Comparison: n8n vs Make vs Zapier
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Capability&lt;/th&gt;
&lt;th&gt;n8n&lt;/th&gt;
&lt;th&gt;Make&lt;/th&gt;
&lt;th&gt;Zapier&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Open source&lt;/td&gt;
&lt;td&gt;Yes (Apache 2.0)&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Self-host&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Code nodes&lt;/td&gt;
&lt;td&gt;JS, Python&lt;/td&gt;
&lt;td&gt;JS&lt;/td&gt;
&lt;td&gt;Limited&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Number of integrations&lt;/td&gt;
&lt;td&gt;400+&lt;/td&gt;
&lt;td&gt;1000+&lt;/td&gt;
&lt;td&gt;6000+&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Per-task pricing&lt;/td&gt;
&lt;td&gt;Free self-host / $24/mo cloud&lt;/td&gt;
&lt;td&gt;$9-$299/mo&lt;/td&gt;
&lt;td&gt;$19.99-$599/mo&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Best for&lt;/td&gt;
&lt;td&gt;Engineering teams&lt;/td&gt;
&lt;td&gt;Mid-market ops teams&lt;/td&gt;
&lt;td&gt;Non-technical teams&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;n8n documentation: &lt;a href="https://docs.n8n.io/" rel="noopener noreferrer"&gt;https://docs.n8n.io/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;n8n workflow templates: &lt;a href="https://n8n.io/workflows/" rel="noopener noreferrer"&gt;https://n8n.io/workflows/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;n8n source code: &lt;a href="https://github.com/n8n-io/n8n" rel="noopener noreferrer"&gt;https://github.com/n8n-io/n8n&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;LangChain + n8n integration: &lt;a href="https://docs.n8n.io/integrations/builtin/cluster-nodes/root-nodes/langchain/" rel="noopener noreferrer"&gt;https://docs.n8n.io/integrations/builtin/cluster-nodes/root-nodes/langchain/&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Acknowledgments
&lt;/h2&gt;

&lt;p&gt;This article summarizes documented patterns; always consult current n8n documentation for version-specific behavior. For legal-workflow-specific patterns, consult your state's ethics opinions on technology competence.&lt;/p&gt;

&lt;p&gt;Dillon Deutsch built CourtGPT.ai using these n8n patterns. &lt;a href="https://courtgpt.ai" rel="noopener noreferrer"&gt;https://courtgpt.ai&lt;/a&gt;&lt;/p&gt;

</description>
      <category>n8n</category>
      <category>legaltech</category>
      <category>automation</category>
      <category>rag</category>
    </item>
    <item>
      <title>Production RAG at Scale: Architecture Patterns for 1M+ Documents</title>
      <dc:creator>CourtGPT</dc:creator>
      <pubDate>Sun, 26 Jul 2026 01:48:04 +0000</pubDate>
      <link>https://dev.to/courtgpt/building-production-rag-systems-lessons-from-67m-legal-records-54k5</link>
      <guid>https://dev.to/courtgpt/building-production-rag-systems-lessons-from-67m-legal-records-54k5</guid>
      <description>&lt;h1&gt;
  
  
  Production RAG at Scale: Architecture Patterns for 1M+ Documents
&lt;/h1&gt;

&lt;p&gt;Retrieval-augmented generation (RAG) systems for production require different design than demo-scale systems. This article summarizes documented patterns for systems indexing 1M+ documents, drawing on published engineering reports and community best practices.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why scale changes everything
&lt;/h2&gt;

&lt;p&gt;At demo scale (&amp;lt;100 docs), basic embedding + cosine similarity works. At production scale (1M+ docs), several patterns emerge:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Embedding model selection matters more&lt;/li&gt;
&lt;li&gt;Retrieval latency grows non-linearly with corpus&lt;/li&gt;
&lt;li&gt;Hallucination rates grow with irrelevant retrievals&lt;/li&gt;
&lt;li&gt;Cost per query can become prohibitive&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This article focuses on architectures that handle these constraints.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tiered retrieval architecture
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Query
  -&amp;gt; Query rewriting (LLM)
  -&amp;gt; Embedding model
  -&amp;gt; Tier 1: BM25 keyword search (top-100)
  -&amp;gt; Tier 2: Vector similarity (top-100)
  -&amp;gt; Reciprocal rank fusion (RRF) -&amp;gt; top-50
  -&amp;gt; Cross-encoder rerank -&amp;gt; top-10
  -&amp;gt; LLM with context window
  -&amp;gt; Streaming response
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Source: Microsoft's RAG research, "Retrieval Augmented Generation for Large Language Models: A Survey" (Gao et al., 2024). Pre-print at &lt;a href="https://arxiv.org/abs/2312.10997" rel="noopener noreferrer"&gt;https://arxiv.org/abs/2312.10997&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Embedding model selection at scale
&lt;/h2&gt;

&lt;p&gt;For 1M+ docs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;text-embedding-3-large&lt;/strong&gt; (OpenAI): 1536-3072 dims, $0.13/M tokens. Mature API. Best for general use.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;voyage-large-2&lt;/strong&gt; (Voyage AI): Competitive on retrieval benchmarks, especially legal/medical domains.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;cohere-embed-v3&lt;/strong&gt; (Cohere): 1024 dims, supports compression. Good for cost-sensitive applications.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;BGE-large-en-v1.5&lt;/strong&gt; (open-source): Self-hostable, no per-token API cost. Slightly lower recall but no API dependency.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Benchmark: MTEB Leaderboard (&lt;a href="https://huggingface.co/spaces/mteb/leaderboard" rel="noopener noreferrer"&gt;https://huggingface.co/spaces/mteb/leaderboard&lt;/a&gt;)&lt;/p&gt;

&lt;h2&gt;
  
  
  Vector database selection
&lt;/h2&gt;

&lt;p&gt;For 1M+ docs with HNSW index:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;pgvector&lt;/strong&gt; (Postgres extension): Self-hostable, transactional, good for hybrid workloads.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pinecone&lt;/strong&gt; (managed): Fast at scale ($70/mo for 1M+ vectors), proprietary.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Weaviate&lt;/strong&gt; (open-source/managed): Hybrid search built-in, BM25 + vector.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Qdrant&lt;/strong&gt; (open-source/managed): Rust-based, low latency.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Milvus&lt;/strong&gt; (open-source): Strong at billion-scale.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For multi-tenant systems, consider per-tenant index partitioning.&lt;/p&gt;

&lt;h2&gt;
  
  
  Hybrid search: BM25 + vectors
&lt;/h2&gt;

&lt;p&gt;For legal/medical text, neither pure keyword nor pure vector retrieval is optimal. Hybrid approach:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Generate BM25 scores (Elasticsearch/OpenSearch)&lt;/li&gt;
&lt;li&gt;Generate vector similarity scores (vector DB)&lt;/li&gt;
&lt;li&gt;Combine via Reciprocal Rank Fusion (RRF)&lt;/li&gt;
&lt;li&gt;Optionally rerank with cross-encoder&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;RRF formula: &lt;code&gt;score(d) = sum(1 / (k + rank_i(d)))&lt;/code&gt; for each ranking.&lt;/p&gt;

&lt;p&gt;Source: Cormack et al. (2009), "Reciprocal Rank Fusion outperforms Condorcet and individual Rank Learning Methods."&lt;/p&gt;

&lt;h2&gt;
  
  
  Chunking strategy
&lt;/h2&gt;

&lt;p&gt;For 1M+ docs at scale:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Recursive character chunking&lt;/strong&gt;: ~95% recall at reasonable compute cost&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Semantic chunking&lt;/strong&gt;: ~98% recall, but 5-10x index time&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Structural chunking&lt;/strong&gt; (legal: Article/Section): Best for legal texts&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Recommended chunk size: 256-512 tokens, 10-20% overlap.&lt;/p&gt;

&lt;p&gt;Source: LangChain text splitters documentation (&lt;a href="https://python.langchain.com/docs/modules/data_connection/document_transformers/" rel="noopener noreferrer"&gt;https://python.langchain.com/docs/modules/data_connection/document_transformers/&lt;/a&gt;)&lt;/p&gt;

&lt;h2&gt;
  
  
  Latency budget
&lt;/h2&gt;

&lt;p&gt;For interactive chat systems:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Target p95 &amp;lt;2s end-to-end&lt;/li&gt;
&lt;li&gt;Typical breakdown:

&lt;ul&gt;
&lt;li&gt;Query rewrite: 200-400ms&lt;/li&gt;
&lt;li&gt;Embedding: 100-300ms&lt;/li&gt;
&lt;li&gt;BM25 + vector search: 100-200ms&lt;/li&gt;
&lt;li&gt;Rerank: 200-500ms&lt;/li&gt;
&lt;li&gt;LLM first-token: 500-1500ms&lt;/li&gt;
&lt;li&gt;Streaming: ongoing&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If p95 &amp;gt;3s, user retention drops materially.&lt;/p&gt;

&lt;h2&gt;
  
  
  Hallucination mitigation
&lt;/h2&gt;

&lt;p&gt;For production legal RAG:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Constrain LLM: "Cite only documents provided. Cite as [N] in response."&lt;/li&gt;
&lt;li&gt;Verify citations post-hoc: Match each cited [N] to actual source.&lt;/li&gt;
&lt;li&gt;Track hallucination rate weekly: Should be &amp;lt;2% in legal contexts.&lt;/li&gt;
&lt;li&gt;Use eval framework (ragas, TruLens, DeepEval).&lt;/li&gt;
&lt;li&gt;Have human review for high-stakes queries.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Source: Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools (Magesh et al., 2024). Working paper at &lt;a href="https://reglab.stanford.edu/publications/hallucination-free-assessing-reliability-leading-ai-legal-research-tools" rel="noopener noreferrer"&gt;https://reglab.stanford.edu/publications/hallucination-free-assessing-reliability-leading-ai-legal-research-tools&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Cost optimization
&lt;/h2&gt;

&lt;p&gt;Top 5 levers (estimated monthly savings at 100K queries/mo):&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Query cache (Redis): 60% hit rate -&amp;gt; -30% LLM cost&lt;/li&gt;
&lt;li&gt;Mix models (Haiku for simple, Sonnet for complex): -50% LLM cost&lt;/li&gt;
&lt;li&gt;Reduce retrieval R from 10 to 5: -40% token cost&lt;/li&gt;
&lt;li&gt;Use smaller embedding model: -25% embedding cost&lt;/li&gt;
&lt;li&gt;Batch async queries: -10% system cost&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Realistic monthly cost at 100K queries: $300-800.&lt;/p&gt;

&lt;h2&gt;
  
  
  Eval framework
&lt;/h2&gt;

&lt;p&gt;Track weekly:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Retrieval precision@K (% of retrieved docs that are relevant)&lt;/li&gt;
&lt;li&gt;Retrieval recall@K (% of relevant docs that are retrieved)&lt;/li&gt;
&lt;li&gt;Faithfulness (does answer stick to retrieved context)&lt;/li&gt;
&lt;li&gt;Answer relevance (is answer topical)&lt;/li&gt;
&lt;li&gt;Citation accuracy (% of citations that resolve to real sources)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Tools:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;ragas (Python): &lt;a href="https://docs.ragas.io/" rel="noopener noreferrer"&gt;https://docs.ragas.io/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;TruLens: &lt;a href="https://www.trulens.org/" rel="noopener noreferrer"&gt;https://www.trulens.org/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;DeepEval: &lt;a href="https://docs.confident-ai.com/" rel="noopener noreferrer"&gt;https://docs.confident-ai.com/&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Operational considerations
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Monitor drift: Re-evaluate weekly to catch regressions&lt;/li&gt;
&lt;li&gt;Version-pin models: Embedding and LLM model versions matter&lt;/li&gt;
&lt;li&gt;Cache invalidation: When corpus changes, invalidate stale cache&lt;/li&gt;
&lt;li&gt;Multi-region: For global latency, deploy in multiple regions&lt;/li&gt;
&lt;li&gt;Cost alerts: Set budget alerts at the LLM provider level&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Gao et al. (2024). Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv:2312.10997&lt;/li&gt;
&lt;li&gt;Magesh et al. (2024). Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. Stanford RegLab.&lt;/li&gt;
&lt;li&gt;Cormack et al. (2009). Reciprocal Rank Fusion outperforms Condorcet and individual Rank Learning Methods.&lt;/li&gt;
&lt;li&gt;MTEB Leaderboard: &lt;a href="https://huggingface.co/spaces/mteb/leaderboard" rel="noopener noreferrer"&gt;https://huggingface.co/spaces/mteb/leaderboard&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;LangChain documentation: &lt;a href="https://python.langchain.com/" rel="noopener noreferrer"&gt;https://python.langchain.com/&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Acknowledgments
&lt;/h2&gt;

&lt;p&gt;This article summarizes documented engineering practices. For specific implementations, consult official documentation and benchmark on your own data.&lt;/p&gt;

&lt;p&gt;Dillon Deutsch has built production RAG systems serving users across all 50 US states. &lt;a href="https://courtgpt.ai" rel="noopener noreferrer"&gt;https://courtgpt.ai&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>rag</category>
      <category>production</category>
      <category>engineering</category>
    </item>
  </channel>
</rss>
