<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Joung Park</title>
    <description>The latest articles on DEV Community by Joung Park (@joungpark).</description>
    <link>https://dev.to/joungpark</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F716955%2Fffdd6810-23b5-413d-94c2-d9830183b088.jpeg</url>
      <title>DEV Community: Joung Park</title>
      <link>https://dev.to/joungpark</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/joungpark"/>
    <language>en</language>
    <item>
      <title>Monorepo or Polyrepo?</title>
      <dc:creator>Joung Park</dc:creator>
      <pubDate>Wed, 07 Oct 2026 05:52:16 +0000</pubDate>
      <link>https://dev.to/joungpark/monorepo-or-polyrepo-1gib</link>
      <guid>https://dev.to/joungpark/monorepo-or-polyrepo-1gib</guid>
      <description>&lt;p&gt;After deciding to use microservices, the next question was:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How should I organise the codebase?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I chose a &lt;strong&gt;monorepo&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Second-Memory has several applications and services:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;apps/
  web/
  mobile/

services/
  memory-service/
  ask-service/
  embedding-worker/

packages/
  api-types/
  ui-components/
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;They are independently deployable, but they are all part of the same product.&lt;/p&gt;

&lt;h2&gt;
  
  
  Monorepo vs Polyrepo
&lt;/h2&gt;

&lt;p&gt;There is no universally better choice.&lt;/p&gt;

&lt;p&gt;The trade-offs become clearer when looking at common development tasks:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Area&lt;/th&gt;
&lt;th&gt;Monorepo&lt;/th&gt;
&lt;th&gt;Polyrepo&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Share code&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Easy&lt;/strong&gt; — reusable code can live in shared packages&lt;/td&gt;
&lt;td&gt;Harder — publish packages or duplicate code&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Share types&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Easy&lt;/strong&gt; — TypeScript types can be shared directly&lt;/td&gt;
&lt;td&gt;Harder — types need separate versioning&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Cross-project changes&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Easy&lt;/strong&gt; — related changes can be made in one PR&lt;/td&gt;
&lt;td&gt;Harder — changes may require multiple PRs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Dependency management&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Easy&lt;/strong&gt; — dependencies can be managed consistently&lt;/td&gt;
&lt;td&gt;More complex — each repo manages its own dependencies&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Local development&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Easy&lt;/strong&gt; — one repository contains the whole system&lt;/td&gt;
&lt;td&gt;More setup — multiple repositories need to be configured&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;CI/CD&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;More complex — need to detect affected projects&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Easy&lt;/strong&gt; — each repository naturally has its own pipeline&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Service isolation&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Harder — boundaries need to be enforced within the repository&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Easy&lt;/strong&gt; — repositories provide natural isolation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Access control&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Harder — repository access generally covers the whole codebase&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Easy&lt;/strong&gt; — permissions can be managed per repository&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Independent releases&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Possible — CI/CD can deploy projects independently&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Easy&lt;/strong&gt; — each repository has its own release lifecycle&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Neither approach is inherently better.&lt;/p&gt;

&lt;p&gt;For Second-Memory, I valued &lt;strong&gt;easy development across the whole system&lt;/strong&gt; more than repository-level isolation.&lt;/p&gt;

&lt;h2&gt;
  
  
  When to Choose Which
&lt;/h2&gt;

&lt;p&gt;A monorepo can be a good fit when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Small team or solo developer&lt;/strong&gt; — one repository is easier to manage than several.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Projects share contracts or code&lt;/strong&gt; — changes can be made and tested together.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cross-project changes are common&lt;/strong&gt; — one PR can update multiple applications or services.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The whole system is closely related&lt;/strong&gt; — one repository makes it easier to understand and work across the product.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Polyrepo can be a better fit when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Large or independent teams&lt;/strong&gt; need clear ownership boundaries.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Strong repository isolation&lt;/strong&gt; is important.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Projects have different release cycles&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Access control&lt;/strong&gt; needs to differ between projects.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For Second-Memory, being a &lt;strong&gt;sole developer&lt;/strong&gt; was one of the main reasons I chose a monorepo.&lt;/p&gt;

&lt;p&gt;Managing several repositories would add overhead without giving me much benefit. I didn't need repository-level ownership boundaries between teams.&lt;/p&gt;

&lt;p&gt;Instead, I wanted to make it easy to work across the web app, mobile app, backend services, and shared contracts.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Should Be Shared?
&lt;/h2&gt;

&lt;p&gt;A monorepo makes sharing easy.&lt;/p&gt;

&lt;p&gt;But &lt;strong&gt;easy to share doesn't mean everything should be shared&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;I wanted to keep the shared surface area small.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Shared Part&lt;/th&gt;
&lt;th&gt;Used By&lt;/th&gt;
&lt;th&gt;Why Share It?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;API types&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Web, Mobile, Backend&lt;/td&gt;
&lt;td&gt;Keep API contracts consistent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Web-specific code&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Web&lt;/td&gt;
&lt;td&gt;Keep it inside &lt;code&gt;apps/web&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Mobile-specific code&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Mobile&lt;/td&gt;
&lt;td&gt;Keep it inside &lt;code&gt;apps/mobile&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Service domain logic&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Individual service&lt;/td&gt;
&lt;td&gt;Keep ownership within the service&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Service implementation&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Individual service&lt;/td&gt;
&lt;td&gt;Preserve service boundaries&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;So instead of creating shared packages for everything, I only create a package when there is a genuine reason to share something.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;apps/
  web/
  mobile/
services/
  memory-service/
  ask-service/
  embedding-worker/
packages/
  api-types/
  ui-components/
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The API types are shared because they represent a contract between different parts of the system.&lt;/p&gt;

&lt;p&gt;But a Web-only component stays in &lt;code&gt;apps/web&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;A Mobile-only utility stays in &lt;code&gt;apps/mobile&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;And Memory Service domain logic stays inside &lt;code&gt;memory-service&lt;/code&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Share the contract, not the ownership
&lt;/h3&gt;

&lt;p&gt;The fact that services live in the same repository doesn't mean they should share their internal implementation.&lt;/p&gt;

&lt;p&gt;The API contract can be shared:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TB
    T[packages/api-types]

    W[Web] --&amp;gt; T
    M[Mobile] --&amp;gt; T
    B[Backend Services] --&amp;gt; T&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;while the services remain independently owned:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TB
    MS[Memory Service]
    AS[Ask Service]

    MS --&amp;gt; MD[Memory Domain]
    AS --&amp;gt; AD[AI Interaction]

    MD --&amp;gt; DB1[(Memory DB)]
    AD --&amp;gt; DB2[(Ask Data)]&lt;/code&gt;&lt;/pre&gt;



&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Share contracts where it makes sense, but don't share service ownership.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Keeping the Shared Surface Small
&lt;/h2&gt;

&lt;p&gt;There is another practical reason for limiting shared code.&lt;/p&gt;

&lt;p&gt;The more projects depend on a shared package, the larger the potential impact of a change.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TB
    S[packages/shared]
    S --&amp;gt; W[Web]
    S --&amp;gt; M[Mobile]
    S --&amp;gt; ME[Memory]
    S --&amp;gt; A[Ask]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;A change to &lt;code&gt;shared&lt;/code&gt; may require all projects to be tested.&lt;/p&gt;

&lt;p&gt;But a change here:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    C[Change in apps/web] --&amp;gt; W[Web]
    W --&amp;gt; CI[Web CI/CD]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;only affects Web.&lt;/p&gt;

&lt;p&gt;This has implications for both &lt;strong&gt;CI/CD&lt;/strong&gt; and &lt;strong&gt;application dependencies&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;I don't want a large &lt;code&gt;shared&lt;/code&gt; package that contains unrelated functionality used by almost every project.&lt;/p&gt;

&lt;p&gt;Instead, I prefer small, focused shared packages with clear reasons for existing.&lt;/p&gt;

&lt;h2&gt;
  
  
  The CI/CD Trade-off
&lt;/h2&gt;

&lt;p&gt;CI/CD is one area where polyrepo can be simpler.&lt;/p&gt;

&lt;p&gt;With polyrepo, a change to &lt;code&gt;memory-service&lt;/code&gt; naturally belongs to the &lt;code&gt;memory-service&lt;/code&gt; repository.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TB
    R[memory-service repo]
    R --&amp;gt; CI[CI/CD]
    CI --&amp;gt; T[Test]
    CI --&amp;gt; B[Build]
    CI --&amp;gt; D[Deploy]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;In a monorepo, multiple deployable units live in the same repository.&lt;/p&gt;

&lt;p&gt;A change to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;services/memory-service/
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;shouldn't necessarily rebuild and deploy everything.&lt;/p&gt;

&lt;p&gt;The pipeline needs to understand:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;what changed&lt;/li&gt;
&lt;li&gt;what depends on it&lt;/li&gt;
&lt;li&gt;what needs to be tested&lt;/li&gt;
&lt;li&gt;what needs to be built&lt;/li&gt;
&lt;li&gt;what needs to be deployed&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This becomes particularly important when shared packages are involved.&lt;/p&gt;

&lt;p&gt;A change to &lt;code&gt;ui-components&lt;/code&gt; may affect only web and mobile projects, while a change isolated to &lt;code&gt;services/embedding-worker/&lt;/code&gt; should ideally affect only the relevant pipeline.&lt;/p&gt;

&lt;p&gt;The repository structure itself is simple, but the build and deployment system needs to become more intelligent as the project grows.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Monorepo for Second-Memory?
&lt;/h2&gt;

&lt;p&gt;The main benefits for me were:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Easy&lt;/strong&gt; sharing of API types&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Easy&lt;/strong&gt; cross-project changes&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Easy&lt;/strong&gt; local development&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Easy&lt;/strong&gt; dependency management&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Easy&lt;/strong&gt; visibility of the whole product&lt;/li&gt;
&lt;li&gt;Less overhead for a &lt;strong&gt;solo developer&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The trade-offs were:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;More complex&lt;/strong&gt; CI/CD&lt;/li&gt;
&lt;li&gt;More tooling needed as the repository grows&lt;/li&gt;
&lt;li&gt;Less repository-level isolation&lt;/li&gt;
&lt;li&gt;The need to carefully control shared dependencies&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For Second-Memory, I was willing to accept these trade-offs.&lt;/p&gt;

&lt;p&gt;The goal wasn't to share as much code as possible.&lt;/p&gt;

&lt;p&gt;The goal was to get the convenience of a monorepo while keeping &lt;strong&gt;ownership and dependencies clear&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That led to the next question:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How do I manage builds, dependencies, task execution, and caching across all these projects?&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>architecture</category>
      <category>microservices</category>
      <category>monorepo</category>
      <category>polyrepo</category>
    </item>
    <item>
      <title>Microservices or Monolith?</title>
      <dc:creator>Joung Park</dc:creator>
      <pubDate>Wed, 07 Oct 2026 04:18:28 +0000</pubDate>
      <link>https://dev.to/joungpark/why-microservices-instead-of-a-monolith-3ld3</link>
      <guid>https://dev.to/joungpark/why-microservices-instead-of-a-monolith-3ld3</guid>
      <description>&lt;p&gt;The initial architecture had three main backend responsibilities:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Memory Service&lt;/li&gt;
&lt;li&gt;Ask Service&lt;/li&gt;
&lt;li&gt;Embedding Worker&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But there was another architectural decision behind this:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why make them separate services instead of building one backend?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A monolith would have been simpler.&lt;/p&gt;

&lt;p&gt;For a small application, it would probably have been enough.&lt;/p&gt;

&lt;p&gt;I still chose a microservice architecture for Second-Memory for three main reasons.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Clear boundaries
&lt;/h2&gt;

&lt;p&gt;Memory and Ask have different responsibilities.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;Memory Service&lt;/strong&gt; owns the memory domain:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;create and manage memories&lt;/li&gt;
&lt;li&gt;persist memory data&lt;/li&gt;
&lt;li&gt;generate/retrieve relevant memories&lt;/li&gt;
&lt;li&gt;manage embeddings&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The &lt;strong&gt;Ask Service&lt;/strong&gt; owns the AI interaction:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;receive questions&lt;/li&gt;
&lt;li&gt;retrieve relevant context&lt;/li&gt;
&lt;li&gt;construct prompts&lt;/li&gt;
&lt;li&gt;call the LLM&lt;/li&gt;
&lt;li&gt;return the answer&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;They work together, but they represent different responsibilities.&lt;/p&gt;

&lt;p&gt;I wanted those boundaries to exist at the service level rather than only as modules inside one application.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    A[API Gateway] --&amp;gt; B[Memory Service]
    A --&amp;gt; C[Ask Service]
    B -.-&amp;gt; D[Memory + Vector]
    C -.-&amp;gt; E[LLM]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The Embedding Worker is separated for a different reason.&lt;/p&gt;

&lt;p&gt;Embedding generation is asynchronous work and doesn't need to be part of the request that creates a memory.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    A[Memory Service] --&amp;gt; B[Outbox]
    B --&amp;gt; C[Queue]
    C --&amp;gt; D[Embedding Worker]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The boundaries therefore follow the responsibilities and execution models of the system.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. A system for learning AI engineering
&lt;/h2&gt;

&lt;p&gt;Second-Memory is also a project for exploring how AI systems are built.&lt;/p&gt;

&lt;p&gt;I didn't want the architecture to stop at:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    A[Application] --&amp;gt; B[LLM API]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The system gave me an opportunity to work with several patterns that are common in production AI systems:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;asynchronous processing&lt;/li&gt;
&lt;li&gt;queues&lt;/li&gt;
&lt;li&gt;workers&lt;/li&gt;
&lt;li&gt;embeddings&lt;/li&gt;
&lt;li&gt;vector search&lt;/li&gt;
&lt;li&gt;retrieval&lt;/li&gt;
&lt;li&gt;service-to-service communication&lt;/li&gt;
&lt;li&gt;LLM integration&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Using separate services makes these boundaries explicit and gives me a way to experiment with each part independently.&lt;/p&gt;

&lt;p&gt;The goal isn't to use microservices simply because they are popular.&lt;/p&gt;

&lt;p&gt;The architecture should help me learn and validate how these pieces work together as a real system.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Production readiness and scaling
&lt;/h2&gt;

&lt;p&gt;I also wanted to design Second-Memory as something that could eventually grow into a production system.&lt;/p&gt;

&lt;p&gt;The different components have different scaling characteristics.&lt;/p&gt;

&lt;p&gt;If memory creation increases, I may need more Memory Service instances.&lt;/p&gt;

&lt;p&gt;If embedding jobs increase, I can scale the workers independently.&lt;/p&gt;

&lt;p&gt;If AI requests become the bottleneck, the Ask Service can scale independently.&lt;/p&gt;

&lt;p&gt;This doesn't mean I need to scale everything independently today.&lt;/p&gt;

&lt;p&gt;It means the architecture doesn't force unrelated workloads into the same deployment unit.&lt;/p&gt;

&lt;h2&gt;
  
  
  What about a monolith?
&lt;/h2&gt;

&lt;p&gt;A monolith would still be a reasonable choice for Second-Memory.&lt;/p&gt;

&lt;p&gt;It would have:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;less infrastructure&lt;/li&gt;
&lt;li&gt;simpler deployment&lt;/li&gt;
&lt;li&gt;simpler local development&lt;/li&gt;
&lt;li&gt;fewer service-to-service calls&lt;/li&gt;
&lt;li&gt;less operational overhead&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So why accept the additional complexity?&lt;/p&gt;

&lt;p&gt;Because my goal wasn't only to build the simplest V1.&lt;/p&gt;

&lt;p&gt;I wanted to build a system with &lt;strong&gt;clear boundaries that could evolve toward production scale&lt;/strong&gt;, while giving me a practical environment to explore AI system architecture.&lt;/p&gt;

&lt;p&gt;The trade-off was intentional.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;I chose microservices not because a monolith couldn't work, but because the boundaries, asynchronous workloads, learning goals, and potential scaling characteristics made the additional complexity worthwhile.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

</description>
      <category>architecture</category>
      <category>microservices</category>
      <category>monolith</category>
    </item>
    <item>
      <title>Why the Outbox Pattern, Queue, and Embedding Worker?</title>
      <dc:creator>Joung Park</dc:creator>
      <pubDate>Wed, 07 Oct 2026 03:47:01 +0000</pubDate>
      <link>https://dev.to/joungpark/why-the-outbox-pattern-queue-and-embedding-worker-2pb5</link>
      <guid>https://dev.to/joungpark/why-the-outbox-pattern-queue-and-embedding-worker-2pb5</guid>
      <description>&lt;p&gt;The previous post covered why I chose pgvector for semantic search.&lt;/p&gt;

&lt;p&gt;But there was another question:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;When should a memory be embedded?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;At first, it might seem simple:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Create memory
     ▼
Generate embedding
     ▼
Save embedding
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;But this makes memory creation dependent on the embedding process.&lt;/p&gt;

&lt;p&gt;If the embedding provider is slow or unavailable, creating a memory could also fail or become slow.&lt;/p&gt;

&lt;p&gt;I wanted to separate these two operations.&lt;/p&gt;

&lt;p&gt;The architecture became:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Memory Service
      ▼
 PostgreSQL
      ▼
 Outbox Table
      ▼
 Queue / Job Runner
      ▼
 Embedding Worker
      ▼
 pgvector
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  1. The Outbox Pattern
&lt;/h2&gt;

&lt;p&gt;When a user creates a memory, the Memory Service writes the memory and an outbox event in the same database transaction.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Transaction
┌─────────────────────────────┐
│ INSERT memory               │
│ INSERT outbox event         │
└─────────────────────────────┘
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important part is that they succeed or fail together.&lt;/p&gt;

&lt;p&gt;This avoids a common problem with distributed systems:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Memory saved
     ▼
Embedding event lost
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The outbox gives me a durable record of the work that needs to happen.&lt;/p&gt;

&lt;p&gt;The outbox is not the queue itself.&lt;/p&gt;

&lt;p&gt;It is a reliable bridge between the database transaction and asynchronous processing.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. The Queue / Job Runner
&lt;/h2&gt;

&lt;p&gt;A separate process reads pending outbox events and puts jobs onto a queue.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Outbox
   ▼
Queue
   ▼
Embedding Worker
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This gives the embedding process some useful properties:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;asynchronous processing&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;retries&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;independent scaling&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;failure isolation&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;no need to block the memory API&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The user can save a memory without waiting for the embedding provider.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. The Embedding Worker
&lt;/h2&gt;

&lt;p&gt;The embedding worker has one main responsibility:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;turn memory text into an embedding and store it in pgvector.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TD
    A[Get Memory] --&amp;gt; B[Generate embedding]
    B --&amp;gt; C[Store Vector]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;This keeps embedding-specific logic out of the Memory Service's synchronous request path.&lt;/p&gt;

&lt;p&gt;It also gives me a place to evolve the embedding pipeline later.&lt;/p&gt;

&lt;p&gt;For example, I could change:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;embedding models&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;chunking strategy&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;retry behaviour&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;batch processing&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;embedding dimensions&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;without changing the API used to create a memory.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Why asynchronous?
&lt;/h2&gt;

&lt;p&gt;There is an important consequence of this design:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A newly created memory may not be immediately searchable.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;There can be a small delay between:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Memory created
     │ asynchronous processing
     ▼
Embedding created
     ▼
Memory available for semantic search
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I accepted this trade-off.&lt;/p&gt;

&lt;p&gt;For Second-Memory, immediate persistence is more important than making the embedding operation part of the user's request.&lt;/p&gt;

&lt;p&gt;This is essentially &lt;strong&gt;eventual consistency&lt;/strong&gt; between the memory record and its vector representation.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Failure becomes easier to handle
&lt;/h2&gt;

&lt;p&gt;The asynchronous design also changes how failures work.&lt;/p&gt;

&lt;p&gt;If the embedding provider temporarily fails:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Memory
  ▼
Outbox
  ▼
Queue
  ▼
Embedding Worker - X -&amp;gt; Retry
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The memory itself has already been safely stored.&lt;/p&gt;

&lt;p&gt;The embedding job can be retried without asking the user to submit the memory again.&lt;/p&gt;

&lt;p&gt;That separation was important to me.&lt;/p&gt;

&lt;h2&gt;
  
  
  The decision
&lt;/h2&gt;

&lt;p&gt;The final flow became:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;graph TD
    CM[Create Memory] --&amp;gt; PG

    subgraph PG [PostgreSQL]
            M[Memory]
            OE[Outbox Event]
    end

    OE --&amp;gt; Q[Queue]
    Q --&amp;gt; EW[Embedding Worker]
    EW --&amp;gt; PV[(pgvector in PostgreSQL)]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;Each part has a different responsibility:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Component&lt;/th&gt;
&lt;th&gt;Responsibility&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Memory Service&lt;/td&gt;
&lt;td&gt;Store the memory&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Outbox&lt;/td&gt;
&lt;td&gt;Reliably record the event&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Queue&lt;/td&gt;
&lt;td&gt;Deliver asynchronous work&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Embedding Worker&lt;/td&gt;
&lt;td&gt;Generate embeddings&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;pgvector&lt;/td&gt;
&lt;td&gt;Store and search vectors&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This added some complexity compared with simply generating the embedding inside the API request.&lt;/p&gt;

&lt;p&gt;But it gave me something more important:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;memory creation is no longer tightly coupled to the availability of the embedding pipeline.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That was the trade-off I wanted.&lt;/p&gt;

</description>
      <category>architecture</category>
      <category>backend</category>
      <category>database</category>
      <category>systemdesign</category>
    </item>
    <item>
      <title>Vector database</title>
      <dc:creator>Joung Park</dc:creator>
      <pubDate>Wed, 07 Oct 2026 03:34:17 +0000</pubDate>
      <link>https://dev.to/joungpark/-why-pgvector-2deb</link>
      <guid>https://dev.to/joungpark/-why-pgvector-2deb</guid>
      <description>&lt;p&gt;Second-Memory stores things a user writes over time.&lt;/p&gt;

&lt;p&gt;If a user later asks:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“What did I write about performance problems?”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A normal keyword search isn't always enough.&lt;/p&gt;

&lt;p&gt;The memory might say:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“The database becomes slow when thousands of users query it at the same time.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;There may be no exact keyword match between the question and the memory.&lt;/p&gt;

&lt;p&gt;This is where &lt;strong&gt;semantic search&lt;/strong&gt; becomes useful.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. From text to vectors
&lt;/h2&gt;

&lt;p&gt;I can convert a piece of text into an &lt;strong&gt;embedding&lt;/strong&gt; — a vector representation of its meaning.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"The database becomes slow when thousands of users query it at the same time."
         ↓
[Embedding model]
         ↓
[0.021, -0.183, ...]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The actual vector contains many dimensions, so it isn't meaningful to look at the individual numbers.&lt;/p&gt;

&lt;p&gt;What matters is the relationship between vectors.&lt;/p&gt;

&lt;p&gt;Texts with similar meanings tend to have vectors that are closer together.&lt;/p&gt;

&lt;p&gt;So when a user asks a question, I can:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;Question&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Embedding&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Vector similarity search&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Relevant memories&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This gives Second-Memory a way to retrieve memories based on &lt;strong&gt;meaning&lt;/strong&gt;, rather than just matching words.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Why pgvector?
&lt;/h2&gt;

&lt;p&gt;Once I decided to use embeddings, I needed somewhere to store and search them.&lt;/p&gt;

&lt;p&gt;I considered:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;pgvector&lt;/th&gt;
&lt;th&gt;Pinecone&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Vector search&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Relational data&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Existing PostgreSQL&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Separate infrastructure&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Operational complexity&lt;/td&gt;
&lt;td&gt;Lower&lt;/td&gt;
&lt;td&gt;Higher&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Good fit for V1&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;I chose &lt;strong&gt;pgvector&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The main reason wasn't that pgvector is necessarily better than Pinecone.&lt;/p&gt;

&lt;p&gt;It was that Second-Memory already had a natural place for the vectors:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;the Memory Service's database.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;With pgvector, I could keep the memory and its embedding together.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;erDiagram
    "Memory Service" ||--|| "PostgreSQL + pgvector" : utilizes
    "PostgreSQL + pgvector" ||--|{ Memory : contains

    Memory {
        uuid user_id
        text content
        vector embedding
    }&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;That kept the architecture simple.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Why not introduce a vector database?
&lt;/h2&gt;

&lt;p&gt;A dedicated vector database could make sense at larger scale.&lt;/p&gt;

&lt;p&gt;But introducing one also creates another system to operate and another boundary to manage.&lt;/p&gt;

&lt;p&gt;For V1, I didn't see enough benefit to justify that complexity.&lt;/p&gt;

&lt;p&gt;The Memory Service could own:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;memory data&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;embeddings&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;vector search&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;all within the same data boundary.&lt;/p&gt;

&lt;p&gt;That also reinforced one of the architectural principles from the previous post:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The service that owns the data should own access to it.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The Ask Service doesn't need to know whether semantic search is implemented with pgvector, Pinecone, or something else.&lt;/p&gt;

&lt;p&gt;It simply asks the Memory Service for relevant memories.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. What this gives me
&lt;/h2&gt;

&lt;p&gt;With pgvector, the basic retrieval flow becomes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User question
      ↓
Generate query embedding
      ↓
Memory Service
      ↓
pgvector similarity search
      ↓
Relevant memories
      ↓
Ask Service
      ↓
     LLM
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is the first important piece of the AI architecture.&lt;/p&gt;

&lt;p&gt;The LLM doesn't need to know everything the user has ever written.&lt;/p&gt;

&lt;p&gt;Instead, the system retrieves the memories that are most relevant to the current question and uses those as context.&lt;/p&gt;

&lt;h2&gt;
  
  
  The decision
&lt;/h2&gt;

&lt;p&gt;For Second-Memory V1, I chose:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;PostgreSQL + pgvector&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;because it gave me semantic search without introducing another database.&lt;/p&gt;

&lt;p&gt;It was a pragmatic choice:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;vector search&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;relational data&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;one data boundary&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;less infrastructure&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;simpler development&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>pgvector</category>
      <category>rag</category>
      <category>vectordatabase</category>
      <category>ai</category>
    </item>
    <item>
      <title>The First Architecture Draft</title>
      <dc:creator>Joung Park</dc:creator>
      <pubDate>Wed, 07 Oct 2026 03:24:36 +0000</pubDate>
      <link>https://dev.to/joungpark/the-first-architecture-draft-2ek9</link>
      <guid>https://dev.to/joungpark/the-first-architecture-draft-2ek9</guid>
      <description>&lt;p&gt;I had the product definition, user stories, and backlog.&lt;/p&gt;

&lt;p&gt;Now it was time to design the system.&lt;/p&gt;

&lt;p&gt;I started with the main responsibilities of Second-Memory and came up with my first architecture draft:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart 
    W[Web Client] --&amp;gt; G[API Gateway / BFF]
    M[Mobile Client] --&amp;gt; G

    G --&amp;gt; A[Auth]
    G --&amp;gt; MS[Memory Service]
    G --&amp;gt; AS[Ask Service]

    MS --&amp;gt; D[(Database)]
    MS --&amp;gt; V[(Vector DB)]

    AS --&amp;gt; L[LLM Provider]

    AS --&amp;gt; MS&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The diagram itself is fairly simple, but there were several design principles behind it.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. BFF for the clients
&lt;/h2&gt;

&lt;p&gt;Both the web and mobile clients communicate through the API Gateway / BFF.&lt;/p&gt;

&lt;p&gt;I chose the &lt;strong&gt;Backend for Frontend (BFF)&lt;/strong&gt; pattern so that the clients don't need to communicate directly with the internal services.&lt;/p&gt;

&lt;p&gt;The BFF provides a client-facing API while hiding the internal architecture.&lt;/p&gt;

&lt;p&gt;It also gives me a place to handle things such as authentication, routing, and client-specific requirements.&lt;/p&gt;

&lt;p&gt;This means the clients don't need to know whether a request is handled by the Memory Service, Ask Service, or something else.&lt;/p&gt;

&lt;p&gt;The internal architecture can evolve without making the clients aware of every change.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Each service owns its data
&lt;/h2&gt;

&lt;p&gt;One principle I wanted to follow was:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Don't share databases between services.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A service owns the data that belongs to its responsibility.&lt;/p&gt;

&lt;p&gt;In this architecture, the Memory Service owns both the database and vector database.&lt;/p&gt;

&lt;p&gt;If another service needs memory data, it doesn't connect directly to either database.&lt;/p&gt;

&lt;p&gt;Instead, it calls the Memory Service through an internal API.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart 
    AS[Ask Service] --&amp;gt;|Internal API| MS
    MS[Memory Service] --&amp;gt; D[(Database)]
    MS --&amp;gt; V[(Vector DB)]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;This keeps the data ownership clear and prevents other services from becoming coupled to the Memory Service's storage implementation.&lt;/p&gt;

&lt;p&gt;If the storage implementation changes later, the other service shouldn't need to know.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Separation of concerns
&lt;/h2&gt;

&lt;p&gt;Each component has a specific responsibility.&lt;/p&gt;

&lt;p&gt;The BFF deals with client-facing communication.&lt;/p&gt;

&lt;p&gt;The Memory Service deals with memory.&lt;/p&gt;

&lt;p&gt;The Ask Service deals with the AI interaction.&lt;/p&gt;

&lt;p&gt;The LLM generates the response.&lt;/p&gt;

&lt;p&gt;The databases provide persistence.&lt;/p&gt;

&lt;p&gt;This separation makes the architecture easier to reason about.&lt;/p&gt;

&lt;p&gt;It also means that AI-specific logic doesn't have to spread throughout the rest of the application.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. DDD and bounded contexts
&lt;/h2&gt;

&lt;p&gt;The separation between Memory Service and Ask Service wasn't just about splitting the code into smaller services.&lt;/p&gt;

&lt;p&gt;I was thinking about the different &lt;strong&gt;domains and responsibilities&lt;/strong&gt; involved.&lt;/p&gt;

&lt;p&gt;The Memory Service represents the memory domain:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;creating memories&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;retrieving memories&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;searching memories&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;managing memory data&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The Ask Service represents the AI interaction:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;receiving a question&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;deciding what information is needed&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;requesting relevant memories&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;constructing context&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;calling the LLM&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;generating the answer&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;They are closely related, but they are not the same responsibility.&lt;/p&gt;

&lt;p&gt;This is where the idea of &lt;strong&gt;bounded contexts&lt;/strong&gt; from Domain-Driven Design helped me think about the boundaries.&lt;/p&gt;

&lt;p&gt;The Ask Service uses the Memory Service, but it doesn't become responsible for managing memory itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Encapsulation and API contracts
&lt;/h2&gt;

&lt;p&gt;The service boundary also creates an important form of encapsulation.&lt;/p&gt;

&lt;p&gt;The Ask Service doesn't need to know how the Memory Service stores or searches memories.&lt;/p&gt;

&lt;p&gt;It only needs to know what the Memory Service provides through its API.&lt;/p&gt;

&lt;p&gt;For example, the Ask Service might ask for relevant memories, without knowing whether the Memory Service uses PostgreSQL, pgvector, a different vector store, or some other implementation in the future.&lt;/p&gt;

&lt;p&gt;The API becomes the contract between the services.&lt;/p&gt;

&lt;p&gt;This keeps implementation details inside the service that owns them.&lt;/p&gt;




&lt;p&gt;These principles gave me the initial structure of Second-Memory.&lt;/p&gt;

&lt;p&gt;It wasn't a complicated architecture, and that was intentional.&lt;/p&gt;

&lt;p&gt;I wanted clear boundaries without creating unnecessary complexity.&lt;/p&gt;

&lt;p&gt;With the boundaries in place, I could start looking at the individual technical decisions.&lt;/p&gt;

</description>
      <category>api</category>
      <category>architecture</category>
      <category>backend</category>
      <category>systemdesign</category>
    </item>
    <item>
      <title>Putting on the Product Owner Hat</title>
      <dc:creator>Joung Park</dc:creator>
      <pubDate>Wed, 07 Oct 2026 02:59:25 +0000</pubDate>
      <link>https://dev.to/joungpark/putting-on-the-product-owner-hat-2k1e</link>
      <guid>https://dev.to/joungpark/putting-on-the-product-owner-hat-2k1e</guid>
      <description>&lt;p&gt;I had the idea for Second-Memory.&lt;/p&gt;

&lt;p&gt;But at that point, it was still just an idea. It was too vague to start thinking about architecture or implementation.&lt;/p&gt;

&lt;p&gt;I wanted to develop the idea further before I started building it.&lt;/p&gt;

&lt;p&gt;So I put on the &lt;strong&gt;Product Owner hat&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;I followed the kind of process I had experienced in my professional work.&lt;/p&gt;

&lt;p&gt;I started by defining the initiative. I wrote down what I was trying to achieve and what I wanted the product to be.&lt;/p&gt;

&lt;p&gt;Then I wrote a PRD.&lt;/p&gt;

&lt;p&gt;Writing it made me think through the product in more detail: the problem, the user, the main use cases, the scope, and what I wanted to achieve with the first version.&lt;/p&gt;

&lt;p&gt;From there, I created user stories.&lt;/p&gt;

&lt;p&gt;This helped me look at the product from the user's perspective rather than thinking about how I was going to implement it.&lt;/p&gt;

&lt;p&gt;I then organised the stories into an initial backlog and broke the work down into tasks that I could eventually implement.&lt;/p&gt;

&lt;p&gt;There was another reason I wanted to document all of this.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;I knew I would be using AI tools to help me build the project.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Instead of explaining the project from scratch every time I asked an AI coding tool for help, I could give it the product documents as context.&lt;/p&gt;

&lt;p&gt;The PRD and user stories became a shared reference point — for me as the developer, and for the AI tools I was working with.&lt;/p&gt;

&lt;p&gt;The idea started in my head as something fairly vague.&lt;/p&gt;

&lt;p&gt;By the end of this process, I had something much more concrete:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;an idea → a product definition → user stories → work I could build.&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>product</category>
      <category>softwaredevelopment</category>
      <category>startup</category>
    </item>
    <item>
      <title>AI Project Idea</title>
      <dc:creator>Joung Park</dc:creator>
      <pubDate>Wed, 07 Oct 2026 02:57:54 +0000</pubDate>
      <link>https://dev.to/joungpark/ai-project-idea-4ebn</link>
      <guid>https://dev.to/joungpark/ai-project-idea-4ebn</guid>
      <description>&lt;h1&gt;
  
  
  I Wanted an AI That Would Just Listen
&lt;/h1&gt;

&lt;p&gt;I wanted to learn how to build AI systems.&lt;/p&gt;

&lt;p&gt;So I started looking for a project.&lt;/p&gt;

&lt;p&gt;At the time, I was using ChatGPT like everyone else. I would type something, and it would answer.&lt;/p&gt;

&lt;p&gt;That's what makes a chatbot useful.&lt;/p&gt;

&lt;p&gt;But I started thinking about something different.&lt;/p&gt;

&lt;p&gt;Sometimes I don't want an answer.&lt;/p&gt;

&lt;p&gt;Sometimes I don't want advice.&lt;/p&gt;

&lt;p&gt;I don't want another question.&lt;/p&gt;

&lt;p&gt;I just want to write down what I'm thinking.&lt;/p&gt;

&lt;p&gt;A thought.&lt;/p&gt;

&lt;p&gt;An idea.&lt;/p&gt;

&lt;p&gt;Something that happened today.&lt;/p&gt;

&lt;p&gt;A feeling.&lt;/p&gt;

&lt;p&gt;Something I'm not even sure is important yet.&lt;/p&gt;

&lt;p&gt;Just like talking to myself.&lt;/p&gt;

&lt;h2&gt;
  
  
  What if the AI didn't respond?
&lt;/h2&gt;

&lt;p&gt;Imagine opening an AI application and writing:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"I've been thinking about learning AI more seriously. I don't know where this will take me, but I feel like I should start building something."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;And then...&lt;/p&gt;

&lt;p&gt;Nothing.&lt;/p&gt;

&lt;p&gt;No advice.&lt;/p&gt;

&lt;p&gt;No suggestions.&lt;/p&gt;

&lt;p&gt;No follow-up question.&lt;/p&gt;

&lt;p&gt;It just keeps it.&lt;/p&gt;

&lt;p&gt;That's it.&lt;/p&gt;

&lt;p&gt;Then a few weeks later, I could ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"What have I been thinking about learning AI?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;And now the system responds.&lt;/p&gt;

&lt;p&gt;Not based on a generic conversation.&lt;/p&gt;

&lt;p&gt;Not based on what I happened to tell it five minutes ago.&lt;/p&gt;

&lt;p&gt;But based on the things &lt;strong&gt;I had previously written down&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That was the idea behind Second-Memory.&lt;/p&gt;

&lt;h2&gt;
  
  
  A different kind of AI interaction
&lt;/h2&gt;

&lt;p&gt;Most chatbots are built around a simple interaction:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;You → AI → You → AI → You&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Every message expects a response.&lt;/p&gt;

&lt;p&gt;I wanted something different:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Me → Second-Memory&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Me → Second-Memory&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Me → Second-Memory&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;...&lt;/p&gt;

&lt;p&gt;And only when I have a question:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Me → Second-Memory → Answer based on what I've written&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The system isn't there to constantly talk to me.&lt;/p&gt;

&lt;p&gt;It's there to remember.&lt;/p&gt;

&lt;p&gt;That's where the name came from:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Second-Memory.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A place where I can put the things I might otherwise forget.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why build this?
&lt;/h2&gt;

&lt;p&gt;This started as an AI learning project.&lt;/p&gt;

&lt;p&gt;I wanted to understand what it actually takes to build an AI-powered system rather than simply calling an LLM API.&lt;/p&gt;

&lt;p&gt;And this idea seemed like a good problem to explore.&lt;/p&gt;

&lt;p&gt;How do you store someone's thoughts?&lt;/p&gt;

&lt;p&gt;How do you find the relevant ones later?&lt;/p&gt;

&lt;p&gt;How do you understand that two thoughts written months apart might be related?&lt;/p&gt;

&lt;p&gt;How do you give an AI enough context to answer a question without giving it everything?&lt;/p&gt;

&lt;p&gt;How do you know whether the answer is actually grounded in what the person wrote?&lt;/p&gt;

&lt;p&gt;I didn't know the answers yet.&lt;/p&gt;

&lt;p&gt;I just had the idea.&lt;/p&gt;

&lt;p&gt;So I created a repository.&lt;/p&gt;

&lt;p&gt;The first commit was:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;second-memory&lt;/strong&gt;&lt;em&gt;Your second memory - capture today, remember forever&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;And that's where the project began.&lt;/p&gt;

</description>
      <category>ai</category>
    </item>
  </channel>
</rss>
