<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: AymaneWebDEV</title>
    <description>The latest articles on DEV Community by AymaneWebDEV (@aymanewebdev).</description>
    <link>https://dev.to/aymanewebdev</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4123217%2F9b1020d9-d33f-4132-a16d-a0760b9c3e6c.png</url>
      <title>DEV Community: AymaneWebDEV</title>
      <link>https://dev.to/aymanewebdev</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/aymanewebdev"/>
    <language>en</language>
    <item>
      <title>Building an Autonomous B2B Lead Enrichment &amp; ICP Scoring Agent with n8n and LangChain</title>
      <dc:creator>AymaneWebDEV</dc:creator>
      <pubDate>Mon, 21 Sep 2026 12:35:42 +0000</pubDate>
      <link>https://dev.to/aymanewebdev/building-an-autonomous-b2b-lead-enrichment-icp-scoring-agent-with-n8n-and-langchain-491f</link>
      <guid>https://dev.to/aymanewebdev/building-an-autonomous-b2b-lead-enrichment-icp-scoring-agent-with-n8n-and-langchain-491f</guid>
      <description>&lt;p&gt;Manual lead qualification and CRM updates can become a surprisingly expensive time sink for early-stage SaaS and engineering teams.&lt;/p&gt;

&lt;p&gt;The workflow is usually the same:&lt;/p&gt;

&lt;p&gt;A lead submits a form → someone checks the company → someone researches the role and company size → someone decides whether the lead matches the ICP → someone updates the CRM or alerts sales.&lt;/p&gt;

&lt;p&gt;You can automate most of that pipeline with &lt;strong&gt;n8n&lt;/strong&gt;, a company enrichment API, and an LLM-based evaluation step.&lt;/p&gt;

&lt;p&gt;In this guide, we'll build the architecture behind an autonomous B2B lead qualification workflow, look at the evaluation prompt, and walk through the n8n workflow structure.&lt;/p&gt;

&lt;p&gt;The goal isn't to make the LLM "decide everything." The goal is to give it structured inputs, explicit qualification rules, and a predictable output that downstream automation can consume.&lt;/p&gt;




&lt;h2&gt;
  
  
  1. System Architecture
&lt;/h2&gt;

&lt;p&gt;The pipeline processes an inbound lead through four stages:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[ Inbound Lead Webhook ]
          │
          ▼
[ Domain &amp;amp; Company Enrichment ]
          │
          ▼
[ AI / ICP Evaluation ]
          │
          ▼
[ Alert + CRM Routing ]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  1. Webhook ingestion
&lt;/h3&gt;

&lt;p&gt;The workflow starts when a lead submits a form, signs up for your application, or sends data from another system.&lt;/p&gt;

&lt;p&gt;Example payload:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"email"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"alex@example.com"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Alex Morgan"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"role"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"CTO"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  2. Company enrichment
&lt;/h3&gt;

&lt;p&gt;The workflow extracts the domain from the email address and sends it to an enrichment provider.&lt;/p&gt;

&lt;p&gt;Depending on the provider, you can retrieve information such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Company name&lt;/li&gt;
&lt;li&gt;Industry&lt;/li&gt;
&lt;li&gt;Employee count&lt;/li&gt;
&lt;li&gt;Domain status&lt;/li&gt;
&lt;li&gt;Funding information&lt;/li&gt;
&lt;li&gt;Company description&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This gives the AI evaluator more context than the original form submission alone.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. ICP evaluation
&lt;/h3&gt;

&lt;p&gt;The enriched lead is passed to an LLM agent.&lt;/p&gt;

&lt;p&gt;Instead of asking:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Is this a good lead?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;we provide explicit qualification criteria and require a structured response.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Automated routing
&lt;/h3&gt;

&lt;p&gt;The result can then be used by n8n to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Send a Slack or Telegram notification&lt;/li&gt;
&lt;li&gt;Create or update a CRM record&lt;/li&gt;
&lt;li&gt;Add the lead to a nurture sequence&lt;/li&gt;
&lt;li&gt;Assign a sales priority&lt;/li&gt;
&lt;li&gt;Store the evaluation in PostgreSQL&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  2. Designing the ICP Evaluation Prompt
&lt;/h2&gt;

&lt;p&gt;The most important part of this workflow isn't the AI model itself.&lt;/p&gt;

&lt;p&gt;It's the &lt;strong&gt;contract between the AI and the automation workflow&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;A useful evaluation prompt should define:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;What data the model receives&lt;/li&gt;
&lt;li&gt;What qualifies as a good lead&lt;/li&gt;
&lt;li&gt;What it should do when information is missing&lt;/li&gt;
&lt;li&gt;Exactly what format it must return&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Evaluate the incoming lead against our Ideal Customer Profile (ICP).

Input Data:
- Lead information: {{lead_data}}
- Enriched company data: {{company_data}}

Qualification Criteria:

1. Company size:
   - More than 10 employees is a positive signal.

2. Role:
   - CTO
   - Founder
   - Engineering Lead
   - VP Product

3. Company:
   - Must appear to be a commercial organization.
   - Personal email domains should not receive a Hot classification.

4. Missing information:
   - Do not invent missing company information.
   - If important information cannot be verified, reduce confidence.

Return JSON only:

{
  "score": "Hot | Warm | Unqualified",
  "reasoning": "One sentence explaining the classification.",
  "recommended_action": "Instant Demo | Add to Nurture | Drop"
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important part here is not the word "deterministic."&lt;/p&gt;

&lt;p&gt;LLMs are probabilistic.&lt;/p&gt;

&lt;p&gt;What we're actually doing is &lt;strong&gt;constraining the model's behavior with explicit rules and a structured output contract&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That makes the result much easier for an automation pipeline to consume.&lt;/p&gt;




&lt;h2&gt;
  
  
  3. Connecting the Workflow in n8n
&lt;/h2&gt;

&lt;p&gt;The basic n8n workflow looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Webhook
   │
   ▼
Extract Email Domain
   │
   ▼
Company Enrichment API
   │
   ▼
Merge Lead + Company Data
   │
   ▼
AI ICP Evaluator
   │
   ▼
Parse Structured Output
   │
   ├── Hot ───────► Sales Alert
   │
   ├── Warm ──────► Nurture / CRM
   │
   └── Unqualified ► Archive
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One improvement I'd recommend over a single AI node is adding an explicit &lt;strong&gt;structured-output/parser step&lt;/strong&gt; after the model.&lt;/p&gt;

&lt;p&gt;That gives you a clear boundary between:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;LLM output
     ↓
Validation
     ↓
Automation
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the model produces malformed JSON, the workflow can stop or retry instead of sending unexpected data into your CRM.&lt;/p&gt;




&lt;h2&gt;
  
  
  4. Example n8n Workflow Structure
&lt;/h2&gt;

&lt;p&gt;A simplified workflow can look like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"B2B Lead ICP Scoring"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"nodes"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Lead Inbound Webhook"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"n8n-nodes-base.webhook"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"parameters"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"path"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"lead-submit"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"httpMethod"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"POST"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Company Enrichment"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"n8n-nodes-base.httpRequest"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"parameters"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"url"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://api.example.com/company"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"ICP AI Evaluator"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"@n8n/n8n-nodes-langchain.agent"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"parameters"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"prompt"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Evaluate the lead against the defined ICP rules and return structured JSON."&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Route Lead"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"n8n-nodes-base.switch"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"parameters"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"mode"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"rules"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is intentionally simplified because the exact n8n JSON depends on the versions of the nodes you're using and the credentials configured in your instance.&lt;/p&gt;

&lt;p&gt;For a real deployment, you'll also need to configure:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;AI model credentials&lt;/li&gt;
&lt;li&gt;Enrichment API credentials&lt;/li&gt;
&lt;li&gt;Telegram/Slack credentials&lt;/li&gt;
&lt;li&gt;Output parsing&lt;/li&gt;
&lt;li&gt;Error handling&lt;/li&gt;
&lt;li&gt;CRM integration&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  5. Don't Let the AI Become a Single Point of Failure
&lt;/h2&gt;

&lt;p&gt;One mistake I see in AI automation workflows is putting too much trust in the model.&lt;/p&gt;

&lt;p&gt;For lead scoring, the AI should be one component inside a controlled pipeline.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Raw Lead
   │
   ▼
Schema Validation
   │
   ▼
Company Enrichment
   │
   ▼
Rule-Based Checks
   │
   ▼
AI Evaluation
   │
   ▼
Output Validation
   │
   ▼
Business Routing
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This gives you multiple opportunities to reject bad input before it reaches your CRM.&lt;/p&gt;

&lt;p&gt;You can also keep hard business rules outside the model.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;IF email domain is gmail.com
    → do not classify as Hot

IF employee_count &amp;lt; 10
    → reduce qualification priority

IF role = CTO
    → add positive signal
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The LLM can handle the less structured reasoning, while explicit rules handle things that should never be ambiguous.&lt;/p&gt;




&lt;h2&gt;
  
  
  6. Production Considerations
&lt;/h2&gt;

&lt;p&gt;Once the workflow starts receiving real traffic, there are a few additional concerns.&lt;/p&gt;

&lt;h3&gt;
  
  
  Queue-based execution
&lt;/h3&gt;

&lt;p&gt;For higher webhook volume, n8n can be deployed using queue-based execution with Redis and worker processes.&lt;/p&gt;

&lt;p&gt;This separates incoming webhook handling from workflow execution and can help handle bursts more reliably.&lt;/p&gt;

&lt;h3&gt;
  
  
  Secrets
&lt;/h3&gt;

&lt;p&gt;Don't hardcode API keys inside workflow definitions.&lt;/p&gt;

&lt;p&gt;Store credentials using n8n's credential system or your deployment's secret-management approach.&lt;/p&gt;

&lt;h3&gt;
  
  
  Rate limiting
&lt;/h3&gt;

&lt;p&gt;Your enrichment and LLM providers will have their own limits.&lt;/p&gt;

&lt;p&gt;Add rate limiting and retry strategies so a burst of inbound leads doesn't unexpectedly consume your entire API quota.&lt;/p&gt;

&lt;h3&gt;
  
  
  Error handling
&lt;/h3&gt;

&lt;p&gt;External APIs fail.&lt;/p&gt;

&lt;p&gt;LLM calls can time out.&lt;/p&gt;

&lt;p&gt;Enrichment data can be incomplete.&lt;/p&gt;

&lt;p&gt;A production workflow should define what happens when each of these situations occurs.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Enrichment API fails
        │
        ▼
      Retry
        │
   ┌────┴────┐
   │         │
Success    Failure
   │         │
   ▼         ▼
Continue   Queue / Review
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  7. Where This Gets Interesting
&lt;/h2&gt;

&lt;p&gt;Once the basic lead-scoring workflow works, you can extend it considerably.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;h3&gt;
  
  
  Lead enrichment
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Email
 ↓
Company
 ↓
Employees
 ↓
Industry
 ↓
Funding
 ↓
Technology Stack
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  AI qualification
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Firmographics
+
Role
+
Company Description
+
Inbound Message
        ↓
    ICP Score
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Automated routing
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Hot
 ├── Slack Alert
 ├── CRM Priority = High
 └── Sales Notification

Warm
 └── Nurture Sequence

Unqualified
 └── Archive
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;At that point, n8n becomes less of a simple workflow builder and more of an orchestration layer around your AI services and business logic.&lt;/p&gt;




&lt;h2&gt;
  
  
  8. Final Architecture
&lt;/h2&gt;

&lt;p&gt;The complete architecture can be summarized as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                 ┌─────────────────────┐
                 │   Lead Submission   │
                 └──────────┬──────────┘
                            │
                            ▼
                 ┌─────────────────────┐
                 │ Schema Validation   │
                 └──────────┬──────────┘
                            │
                            ▼
                 ┌─────────────────────┐
                 │ Company Enrichment  │
                 └──────────┬──────────┘
                            │
                            ▼
                 ┌─────────────────────┐
                 │ Rule-Based Checks   │
                 └──────────┬──────────┘
                            │
                            ▼
                 ┌─────────────────────┐
                 │   AI ICP Scoring    │
                 └──────────┬──────────┘
                            │
                            ▼
                 ┌─────────────────────┐
                 │ Output Validation   │
                 └──────────┬──────────┘
                            │
                 ┌──────────┼──────────┐
                 ▼          ▼          ▼
               Hot        Warm      Unqualified
                 │          │          │
                 ▼          ▼          ▼
              Sales      Nurture     Archive
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The main lesson is that you don't need to make the LLM responsible for the entire process.&lt;/p&gt;

&lt;p&gt;Use &lt;strong&gt;code for strict rules&lt;/strong&gt;, &lt;strong&gt;APIs for enrichment&lt;/strong&gt;, &lt;strong&gt;LLMs for flexible evaluation&lt;/strong&gt;, and &lt;strong&gt;n8n for orchestration&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That separation makes the system easier to debug, test, and modify as your ICP changes.&lt;/p&gt;




&lt;h2&gt;
  
  
  Want the Full Automation Vault?
&lt;/h2&gt;

&lt;p&gt;If you want to build beyond lead qualification, I packaged a collection of reusable n8n + AI workflows covering areas such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Lead qualification&lt;/li&gt;
&lt;li&gt;Document RAG&lt;/li&gt;
&lt;li&gt;GitHub issue triaging&lt;/li&gt;
&lt;li&gt;Competitor monitoring&lt;/li&gt;
&lt;li&gt;Customer churn workflows&lt;/li&gt;
&lt;li&gt;AI-powered business automation&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Enterprise n8n &amp;amp; AI Agent Automation Vault
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://nexusbuilds.gumroad.com/l/n8n-ai-automation-vault/EARLYBIRD" rel="noopener noreferrer"&gt;https://nexusbuilds.gumroad.com/l/n8n-ai-automation-vault/EARLYBIRD&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;For the prompt-engineering rules used to structure AI coding and agent workflows, you can also check out the open-source repository:&lt;/p&gt;

&lt;h3&gt;
  
  
  Developer Prompt Vault
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://github.com/AymaneWebDEV/developer-prompt-vault" rel="noopener noreferrer"&gt;https://github.com/AymaneWebDEV/developer-prompt-vault&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If you're building something similar, I'd be interested to know how you're handling &lt;strong&gt;AI scoring vs. rule-based qualification&lt;/strong&gt; in your own automation stack.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>n8nbrightdatachallenge</category>
      <category>automation</category>
      <category>devops</category>
    </item>
    <item>
      <title>How to Build Sub-100ms Streaming AI APIs with Next.js 15 and Supabase SSR</title>
      <dc:creator>AymaneWebDEV</dc:creator>
      <pubDate>Wed, 16 Sep 2026 11:51:57 +0000</pubDate>
      <link>https://dev.to/aymanewebdev/how-to-build-sub-100ms-streaming-ai-apis-with-nextjs-15-and-supabase-ssr-9bc</link>
      <guid>https://dev.to/aymanewebdev/how-to-build-sub-100ms-streaming-ai-apis-with-nextjs-15-and-supabase-ssr-9bc</guid>
      <description>&lt;h1&gt;
  
  
  Building a Secure Streaming AI Endpoint with Next.js App Router
&lt;/h1&gt;

&lt;p&gt;If you're building an AI SaaS or LLM wrapper in 2026, &lt;strong&gt;time-to-first-token (TTFT)&lt;/strong&gt; has a huge impact on how the application feels.&lt;/p&gt;

&lt;p&gt;Waiting several seconds for the entire completion before showing anything makes the UI feel slow, even when the model itself is responding normally.&lt;/p&gt;

&lt;p&gt;Streaming fixes that by sending tokens to the client as they're generated.&lt;/p&gt;

&lt;p&gt;But there's another side to it: implementing streaming carelessly can introduce problems around authentication, API-key exposure, rate limiting, and resource consumption.&lt;/p&gt;

&lt;p&gt;Here's a simple architecture for building a streaming AI endpoint with &lt;strong&gt;Next.js App Router, Supabase authentication, and server-side rate limiting&lt;/strong&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  1. The Streaming Route Handler
&lt;/h2&gt;

&lt;p&gt;Let's start with &lt;code&gt;/api/ai/stream&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The native Web Streams API works well for this use case because we can send chunks to the client as they become available instead of buffering the entire response.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;NextRequest&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;NextResponse&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;next/server&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;createClient&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;@/lib/supabase/server&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;runtime&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;edge&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;POST&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;NextRequest&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="c1"&gt;// 1. Authenticate using Supabase SSR cookies&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;supabase&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;createClient&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="na"&gt;data&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;user&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;supabase&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;auth&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getUser&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;user&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;NextResponse&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;error&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Unauthorized&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;401&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
      &lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;prompt&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

    &lt;span class="c1"&gt;// 2. Create the response stream&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;encoder&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;TextEncoder&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;stream&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;ReadableStream&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
      &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="nf"&gt;start&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;controller&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;chunks&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
          &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Analyzing architecture... &lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
          &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Generating edge pipeline... &lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
          &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Done.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;];&lt;/span&gt;

        &lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;chunk&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="nx"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
          &lt;span class="nx"&gt;controller&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;enqueue&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="nx"&gt;encoder&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
              &lt;span class="s2"&gt;`data: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;chunk&lt;/span&gt; &lt;span class="p"&gt;})}&lt;/span&gt;&lt;span class="s2"&gt;\n\n`&lt;/span&gt;
            &lt;span class="p"&gt;)&lt;/span&gt;
          &lt;span class="p"&gt;);&lt;/span&gt;

          &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Promise&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;resolve&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;setTimeout&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;resolve&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;60&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;

        &lt;span class="nx"&gt;controller&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;enqueue&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
          &lt;span class="nx"&gt;encoder&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;data: [DONE]&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="p"&gt;);&lt;/span&gt;

        &lt;span class="nx"&gt;controller&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;close&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
      &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;});&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Response&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="na"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Content-Type&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;text/event-stream; charset=utf-8&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Cache-Control&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;no-cache, no-transform&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="na"&gt;Connection&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;keep-alive&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;error&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Stream error:&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;error&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;NextResponse&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
      &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;error&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Stream error&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
      &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;500&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The example above uses simulated chunks, but the same stream can be connected to an actual LLM provider.&lt;/p&gt;

&lt;p&gt;The important part is that the server starts sending data immediately rather than waiting for the complete response.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. Keep Your API Keys on the Server
&lt;/h2&gt;

&lt;p&gt;One mistake I see frequently in AI wrappers is putting provider credentials in client-side code.&lt;/p&gt;

&lt;p&gt;Don't do this.&lt;/p&gt;

&lt;p&gt;Your browser should communicate with your application:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Browser
   ↓
/api/ai/stream
   ↓
Authentication
   ↓
Rate limiting
   ↓
LLM provider
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The OpenAI, Anthropic, or other provider API key should remain on the server.&lt;/p&gt;

&lt;p&gt;The client only needs to receive the streamed output.&lt;/p&gt;




&lt;h2&gt;
  
  
  3. Authenticate Before Starting the LLM Request
&lt;/h2&gt;

&lt;p&gt;Streaming makes authentication even more important because an unauthorized user can potentially keep connections open and consume resources.&lt;/p&gt;

&lt;p&gt;The sequence should be:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Request
  ↓
Validate session
  ↓
Check rate limit
  ↓
Validate input
  ↓
Call LLM
  ↓
Stream response
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Don't start the expensive provider request and then perform authorization afterward.&lt;/p&gt;

&lt;p&gt;Reject invalid requests as early as possible.&lt;/p&gt;




&lt;h2&gt;
  
  
  4. Add Rate Limiting
&lt;/h2&gt;

&lt;p&gt;Authentication alone doesn't stop an authenticated user from making thousands of requests.&lt;/p&gt;

&lt;p&gt;For an AI SaaS, this can become an expensive problem very quickly.&lt;/p&gt;

&lt;p&gt;A rate limiter should generally be applied &lt;strong&gt;before&lt;/strong&gt; the LLM request is created.&lt;/p&gt;

&lt;p&gt;Depending on your architecture, you might use Redis or another shared store so limits work across multiple application instances.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;user_123
   │
   ├── request 1 ✓
   ├── request 2 ✓
   ├── request 3 ✓
   └── request 4 → rate limited
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For production systems, consider limiting based on both request frequency and the amount of work being requested.&lt;/p&gt;




&lt;h2&gt;
  
  
  5. Validate the Input
&lt;/h2&gt;

&lt;p&gt;Don't blindly pass arbitrary request bodies to your LLM provider.&lt;/p&gt;

&lt;p&gt;At minimum, validate:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Prompt type&lt;/li&gt;
&lt;li&gt;Maximum prompt length&lt;/li&gt;
&lt;li&gt;Required fields&lt;/li&gt;
&lt;li&gt;Allowed parameters&lt;/li&gt;
&lt;li&gt;Model selection&lt;/li&gt;
&lt;li&gt;Token limits&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This prevents malformed requests from reaching the expensive part of your stack.&lt;/p&gt;




&lt;h2&gt;
  
  
  6. The Architecture
&lt;/h2&gt;

&lt;p&gt;Putting everything together:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                 ┌──────────────────┐
                 │     Browser      │
                 └────────┬─────────┘
                          │
                          ▼
                 ┌──────────────────┐
                 │ /api/ai/stream   │
                 └────────┬─────────┘
                          │
                    Authentication
                          │
                          ▼
                    Rate Limiting
                          │
                          ▼
                    Input Validation
                          │
                          ▼
                 ┌──────────────────┐
                 │   LLM Provider   │
                 └────────┬─────────┘
                          │
                     Token stream
                          │
                          ▼
                 ┌──────────────────┐
                 │   ReadableStream │
                 └────────┬─────────┘
                          │
                          ▼
                       Browser
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The key idea isn't simply "use streaming."&lt;/p&gt;

&lt;p&gt;It's to make streaming part of a properly controlled request pipeline.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Authenticate → rate-limit → validate → call the provider → stream the result.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That gives you a much better foundation for turning an LLM API into an actual SaaS feature rather than exposing a raw model endpoint.&lt;/p&gt;




&lt;h3&gt;
  
  
  One More Thing
&lt;/h3&gt;

&lt;p&gt;Don't assume that &lt;code&gt;runtime = "edge"&lt;/code&gt; automatically means your endpoint will have sub-100ms TTFT.&lt;/p&gt;

&lt;p&gt;The runtime can reduce some latency, but actual TTFT still depends on your authentication path, region, network, provider, model, prompt size, and infrastructure.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Measure it. Don't market the number before you've measured it.&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>nextjs</category>
      <category>supabase</category>
      <category>webdev</category>
      <category>ai</category>
    </item>
    <item>
      <title>How to Stop Cursor from Hallucinating: 5 Production Rules Every AI Engineer Needs</title>
      <dc:creator>AymaneWebDEV</dc:creator>
      <pubDate>Sun, 13 Sep 2026 13:41:37 +0000</pubDate>
      <link>https://dev.to/aymanewebdev/how-to-stop-cursor-from-hallucinating-5-production-rules-every-ai-engineer-needs-18o9</link>
      <guid>https://dev.to/aymanewebdev/how-to-stop-cursor-from-hallucinating-5-production-rules-every-ai-engineer-needs-18o9</guid>
      <description>&lt;p&gt;How to Stop Cursor from Hallucinating: 5 Production Rules Every AI Engineer Needs&lt;/p&gt;

&lt;p&gt;If you use Cursor, Claude Code, or GitHub Copilot on a non-trivial codebase, you've probably encountered the &lt;strong&gt;AI coding drift&lt;/strong&gt; problem.&lt;/p&gt;

&lt;p&gt;The model:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Invents methods that don't exist in your framework version.&lt;/li&gt;
&lt;li&gt;Couples database queries or third-party APIs directly inside HTTP route handlers.&lt;/li&gt;
&lt;li&gt;Generates loose types like &lt;code&gt;any&lt;/code&gt; or &lt;code&gt;object&lt;/code&gt; to bypass TypeScript or Pydantic errors.&lt;/li&gt;
&lt;li&gt;Writes brittle unit tests that mock everything without testing actual boundary failures.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When building production systems, manually fixing AI-generated drift can quickly erase the productivity gains from AI-assisted development.&lt;/p&gt;

&lt;p&gt;The solution isn't simply "better conversational prompting."&lt;/p&gt;

&lt;p&gt;It's &lt;strong&gt;explicit engineering rules that constrain the coding agent before it writes the code.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Here are five rules you can add to your &lt;code&gt;.cursor/rules/&lt;/code&gt; directory.&lt;/p&gt;




&lt;h2&gt;
  
  
  1. Bounded Context &amp;amp; Layer Isolation
&lt;/h2&gt;

&lt;p&gt;Prevent the AI from mixing database access, business logic, and HTTP concerns:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="p"&gt;-&lt;/span&gt; Route handlers MUST only perform request validation and delegate to application services.
&lt;span class="p"&gt;-&lt;/span&gt; Domain logic MUST remain independent of database ORMs and external APIs.
&lt;span class="p"&gt;-&lt;/span&gt; External API clients and third-party SDKs MUST be encapsulated behind dedicated adapters.
&lt;span class="p"&gt;-&lt;/span&gt; Database queries MUST NOT be placed directly inside HTTP route handlers.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This gives the agent explicit boundaries between the presentation, application, domain, and infrastructure layers.&lt;/p&gt;

&lt;p&gt;Without these constraints, an AI agent will often choose the shortest path to a working implementation—even when that implementation creates unnecessary coupling.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. Hermetic Unit Testing
&lt;/h2&gt;

&lt;p&gt;AI-generated tests can look comprehensive while providing very little protection.&lt;/p&gt;

&lt;p&gt;A common pattern is to mock almost every dependency and then assert that the mocked functions were called.&lt;/p&gt;

&lt;p&gt;The test passes, but the actual boundary failure was never tested.&lt;/p&gt;

&lt;p&gt;Give the agent explicit testing constraints:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="p"&gt;-&lt;/span&gt; Unit tests MUST be hermetic: no real network calls and no unintended external filesystem dependencies.
&lt;span class="p"&gt;-&lt;/span&gt; Tests MUST cover valid inputs, invalid inputs, and important boundary conditions.
&lt;span class="p"&gt;-&lt;/span&gt; Avoid mocking internal domain logic.
&lt;span class="p"&gt;-&lt;/span&gt; Mock external boundary adapters where appropriate.
&lt;span class="p"&gt;-&lt;/span&gt; Tests MUST verify observable behavior rather than implementation details.
&lt;span class="p"&gt;-&lt;/span&gt; Every bug fix SHOULD include a regression test when practical.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For example, don't only test that an API call succeeds.&lt;/p&gt;

&lt;p&gt;Also test what happens when the external service:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Times out&lt;/li&gt;
&lt;li&gt;Returns malformed data&lt;/li&gt;
&lt;li&gt;Returns an unexpected status code&lt;/li&gt;
&lt;li&gt;Returns an empty response&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The goal isn't maximum mock coverage.&lt;/p&gt;

&lt;p&gt;It's meaningful behavioral coverage.&lt;/p&gt;




&lt;h2&gt;
  
  
  3. Fail-Fast Input Boundaries
&lt;/h2&gt;

&lt;p&gt;AI-generated applications often assume that incoming data is trustworthy.&lt;/p&gt;

&lt;p&gt;That's particularly dangerous at API boundaries.&lt;/p&gt;

&lt;p&gt;Tell the agent exactly how input should be handled:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="p"&gt;-&lt;/span&gt; All incoming payloads MUST be validated using strict schemas such as Pydantic or Zod.
&lt;span class="p"&gt;-&lt;/span&gt; Avoid implicit type coercion when strict validation is required.
&lt;span class="p"&gt;-&lt;/span&gt; Invalid input MUST be rejected at the application boundary.
&lt;span class="p"&gt;-&lt;/span&gt; Domain-specific failures MUST use typed exceptions rather than generic errors.
&lt;span class="p"&gt;-&lt;/span&gt; Do not pass unvalidated request dictionaries through application layers.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This creates a clear boundary:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;External input → Validation → Application logic → Domain logic&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Instead of allowing malformed data to travel through the entire application before something eventually fails.&lt;/p&gt;




&lt;h2&gt;
  
  
  4. Hallucination &amp;amp; Assumption Defense
&lt;/h2&gt;

&lt;p&gt;One of the most frustrating problems with AI coding assistants is confident guessing.&lt;/p&gt;

&lt;p&gt;A model may generate an import, method, parameter, or dependency that looks perfectly reasonable but doesn't actually exist in your installed version.&lt;/p&gt;

&lt;p&gt;Add rules that explicitly prohibit this behavior:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="p"&gt;-&lt;/span&gt; NEVER assume an API, method, parameter, or configuration option exists without verification.
&lt;span class="p"&gt;-&lt;/span&gt; Prefer APIs already used by the existing codebase when implementing new functionality.
&lt;span class="p"&gt;-&lt;/span&gt; Do not invent dependencies or speculative import paths.
&lt;span class="p"&gt;-&lt;/span&gt; Do not use deprecated APIs when a supported alternative exists.
&lt;span class="p"&gt;-&lt;/span&gt; If an implementation depends on an unverified assumption, explicitly identify the assumption before proceeding.
&lt;span class="p"&gt;-&lt;/span&gt; If the requested approach introduces architectural or security risks, explain the trade-off and propose a safer alternative.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important principle is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The project's installed dependencies and existing code are the source of truth—not the model's memory.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is particularly useful when working with rapidly changing frameworks and libraries.&lt;/p&gt;




&lt;h2&gt;
  
  
  5. Idempotent State Mutations
&lt;/h2&gt;

&lt;p&gt;AI-generated APIs can also overlook what happens when clients retry requests.&lt;/p&gt;

&lt;p&gt;Consider an endpoint that creates an order or processes a payment.&lt;/p&gt;

&lt;p&gt;If the client sends the request, experiences a timeout, and retries it, you don't want the server to process the operation twice.&lt;/p&gt;

&lt;p&gt;For state-changing operations, give the agent explicit constraints:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="p"&gt;-&lt;/span&gt; State-changing endpoints SHOULD support idempotency when duplicate requests could cause unintended side effects.
&lt;span class="p"&gt;-&lt;/span&gt; Financial, order, and payment operations MUST define an idempotency strategy.
&lt;span class="p"&gt;-&lt;/span&gt; Idempotency keys MUST be persisted and associated with the resulting operation.
&lt;span class="p"&gt;-&lt;/span&gt; Concurrent state transitions MUST use appropriate transactional or locking mechanisms.
&lt;span class="p"&gt;-&lt;/span&gt; Do not rely on application-level checks alone when atomic database guarantees are required.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The exact implementation will depend on your database and architecture, but the important thing is that the agent is forced to &lt;strong&gt;consider retry and concurrency behavior&lt;/strong&gt; instead of generating only the happy path.&lt;/p&gt;




&lt;h1&gt;
  
  
  Why These Rules Matter
&lt;/h1&gt;

&lt;p&gt;AI coding assistants are extremely good at generating code.&lt;/p&gt;

&lt;p&gt;But they don't automatically know:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Your architecture&lt;/li&gt;
&lt;li&gt;Your dependency versions&lt;/li&gt;
&lt;li&gt;Your domain boundaries&lt;/li&gt;
&lt;li&gt;Your testing philosophy&lt;/li&gt;
&lt;li&gt;Your security requirements&lt;/li&gt;
&lt;li&gt;Your tolerance for technical debt&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Without explicit constraints, the model tends to optimize for producing code that looks plausible and solves the immediate request.&lt;/p&gt;

&lt;p&gt;That's where coding drift begins.&lt;/p&gt;

&lt;p&gt;Compare:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Implement authentication."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;with:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Use the existing authentication service.&lt;br&gt;
Do not access the database from route handlers.&lt;br&gt;
Validate all external input with the existing schema system.&lt;br&gt;
Do not introduce new dependencies.&lt;br&gt;
Add tests for expired tokens and invalid credentials.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The second instruction gives the agent a much smaller—and more useful—solution space.&lt;/p&gt;




&lt;h1&gt;
  
  
  Start With Constraints, Then Generate Code
&lt;/h1&gt;

&lt;p&gt;You don't need an enormous system prompt containing every possible engineering rule.&lt;/p&gt;

&lt;p&gt;Start with the constraints that matter most to your project:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Architecture boundaries&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Dependency verification&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Strict input and type validation&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Boundary-focused testing&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;State and concurrency safety&lt;/strong&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Then adapt them to your framework and codebase.&lt;/p&gt;

&lt;p&gt;The goal isn't to make your AI coding assistant less autonomous.&lt;/p&gt;

&lt;p&gt;It's to make its autonomy &lt;strong&gt;bounded by engineering constraints&lt;/strong&gt;.&lt;/p&gt;




&lt;h1&gt;
  
  
  Open-Source Developer Prompt Vault
&lt;/h1&gt;

&lt;p&gt;I've collected these types of rules, along with additional prompts for architecture, refactoring, testing, security, and development workflows, in an open-source repository:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Developer Prompt Vault:&lt;/strong&gt;&lt;br&gt;
&lt;a href="https://github.com/AymaneWebDEV/developer-prompt-vault" rel="noopener noreferrer"&gt;https://github.com/AymaneWebDEV/developer-prompt-vault&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The repository contains reusable Markdown rules and templates that you can adapt to your own AI-assisted development workflow.&lt;/p&gt;

&lt;p&gt;If you're building AI-powered applications with FastAPI, I've also put together a separate starter architecture covering authentication, streaming APIs, rate limiting, Docker, and automated testing:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;FastAPI AI Agent Starter Kit:&lt;/strong&gt;&lt;br&gt;
&lt;a href="https://nexusbuilds.gumroad.com/l/fastapi-ai-starter-kit" rel="noopener noreferrer"&gt;https://nexusbuilds.gumroad.com/l/fastapi-ai-starter-kit&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  What Rules Have Helped You?
&lt;/h2&gt;

&lt;p&gt;What constraints have had the biggest impact on your Cursor, Claude Code, or Copilot workflows?&lt;/p&gt;

&lt;p&gt;I'd especially be interested in rules around &lt;strong&gt;architecture, testing, dependency verification, and preventing AI-generated technical debt&lt;/strong&gt;.&lt;/p&gt;

</description>
      <category>cursor</category>
      <category>ai</category>
      <category>programming</category>
      <category>webdev</category>
    </item>
  </channel>
</rss>
