<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Scott Shoemaker</title>
    <description>The latest articles on DEV Community by Scott Shoemaker (@scott_shoemaker_8d10ccbd2).</description>
    <link>https://dev.to/scott_shoemaker_8d10ccbd2</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4013252%2Fad8d22f5-32d4-4a9e-9133-05c112ec6bee.png</url>
      <title>DEV Community: Scott Shoemaker</title>
      <link>https://dev.to/scott_shoemaker_8d10ccbd2</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/scott_shoemaker_8d10ccbd2"/>
    <language>en</language>
    <item>
      <title>Crossing the Deployment Bridge: Cloud Run, CORS, and Preparing for User Testing</title>
      <dc:creator>Scott Shoemaker</dc:creator>
      <pubDate>Mon, 05 Oct 2026 16:17:58 +0000</pubDate>
      <link>https://dev.to/scott_shoemaker_8d10ccbd2/crossing-the-deployment-bridge-cloud-run-cors-and-preparing-for-user-testing-ook</link>
      <guid>https://dev.to/scott_shoemaker_8d10ccbd2/crossing-the-deployment-bridge-cloud-run-cors-and-preparing-for-user-testing-ook</guid>
      <description>&lt;p&gt;Getting a full-stack application running flawlessly on your local machine is a great feeling. Getting that same application to run flawlessly on the live internet is an entirely different beast.&lt;/p&gt;

&lt;p&gt;This past week on MedReach AI, my teammate Collin and I tackled the deployment transition, moving our testing from the local emulator to a live production environment. We successfully deployed our FastAPI backend to Google Cloud Run and our React frontend to Firebase Hosting, but the journey wasn't without its hurdles.&lt;/p&gt;

&lt;p&gt;Here is a look at the technical challenges we solved this week, and how we are preparing for our upcoming user testing phase.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The CORS Bouncer and the Cloud Run Fix&lt;/strong&gt;&lt;br&gt;
Right after our initial Firebase deployment, our frontend dashboard threw a wall of "Failed to fetch" errors under our executive metrics. The backend was alive, the frontend was live, but they were refusing to talk to each other.&lt;/p&gt;

&lt;p&gt;The culprit was a classic CORS (Cross-Origin Resource Sharing) block. Our Google Cloud Run backend was acting like a strict bouncer, rejecting requests from our brand new Firebase URL because it wasn't on the approved list.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The Solution&lt;/em&gt;: We had to carefully strip out our local development overrides (like bypassing offline Firestore configurations) and update our main.py file to explicitly whitelist our Firebase Hosting origin in the FastAPI CORSMiddleware. By pushing these changes through our GitHub pipeline, Cloud Run automatically pulled the latest revision. Within minutes, the CORS block was lifted, and our live executive metrics populated perfectly.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F22fvycljqqil8ckrby4y.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F22fvycljqqil8ckrby4y.png" alt=" " width="800" height="425"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Redesigning the Data Health Score UX&lt;/strong&gt;&lt;br&gt;
Once the data pipeline was connected, we realized we had a critical UX flaw on the Data Review screen.&lt;/p&gt;

&lt;p&gt;Our system grades datasets across different categories, but the UI displayed the points as fractions (e.g., 20/20 or 21.3/25). At a glance, a perfect 25/25 score looked exactly like "25 unresolved errors out of 25," creating instant panic for the user. Furthermore, our PII/PHI (Personally Identifiable Information) scanner was acting purely as an advisory warning rather than deducting points from the overall data health score—a major oversight for a healthcare data platform.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The Solution:&lt;/em&gt; We overhauled the scoring logic and the UI. We rebalanced the 100-point total across five distinct 20-point categories, finally giving PII/PHI its own dedicated scoring bucket. If sensitive data leaks are detected, the dataset is actively penalized. On the frontend, we appended "pts" to the fractions (e.g., 20/20 pts) and grouped the record-level privacy detections under a dedicated tab. This small change completely removed the ambiguity between "points earned" and "flags detected."&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fk149h1scv8qk88y98zh3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fk149h1scv8qk88y98zh3.png" alt=" " width="800" height="423"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Sprint Week 2: The "Think-Aloud" User Testing&lt;/strong&gt;&lt;br&gt;
With the production environment stable and the UI refined, this week is all about stepping back and letting fresh eyes tear it apart.&lt;/p&gt;

&lt;p&gt;We are officially entering our user testing sprint. I have scheduled three Think-Aloud testing sessions with Melissa, Alex, and Aaron for October 7th through the 9th. Because they have not spent the last several weeks staring at this specific codebase, they are the perfect candidates to expose any assumptions we have made about the platform's intuitive flow.&lt;/p&gt;

&lt;p&gt;During these sessions, I will be observing how they navigate the CSV upload process and whether the newly updated Data Review screen actually makes sense to a first-time user. Meanwhile, Collin is diving deep into end-to-end integration testing to ensure our backend handles payload mismatches without throwing 500-level errors.&lt;/p&gt;

&lt;p&gt;Building software in a vacuum is dangerous. It’s time to see how MedReach AI holds up in the wild. I'll share the results of our testing sessions next week!&lt;/p&gt;

</description>
      <category>ai</category>
      <category>beginners</category>
      <category>testing</category>
      <category>development</category>
    </item>
    <item>
      <title>Week 1: Stabilizing the Handshake — Overcoming E2E Integration Roadblocks</title>
      <dc:creator>Scott Shoemaker</dc:creator>
      <pubDate>Tue, 29 Sep 2026 16:43:47 +0000</pubDate>
      <link>https://dev.to/scott_shoemaker_8d10ccbd2/week-1-stabilizing-the-handshake-overcoming-e2e-integration-roadblocks-2cif</link>
      <guid>https://dev.to/scott_shoemaker_8d10ccbd2/week-1-stabilizing-the-handshake-overcoming-e2e-integration-roadblocks-2cif</guid>
      <description>&lt;p&gt;Welcome to the final phase of development for MedReachAI. As we transition into the Software Integration phase, my partner, Collin, and I are shifting our focus from building core features to hardening our application for real-world usability, security, and stability.&lt;/p&gt;

&lt;p&gt;This week, we tackled a massive milestone: getting our React frontend and FastAPI backend to talk to each other flawlessly under heavy data loads. It hasn't been without its roadblocks, but resolving those bottlenecks is exactly what this phase is about.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Problem: Memory Crashes and the Integration Gap&lt;/strong&gt;&lt;br&gt;
MedReachAI is designed to ingest, clean, and analyze massive healthcare provider datasets. Previously, when we attempted to push large CSV files directly through our web server, we hit immediate bottlenecks. Cloud Run payload limits and Out of Memory (OOM) container crashes stopped our pipeline in its tracks.&lt;/p&gt;

&lt;p&gt;To solve this, we pivoted to a Direct-to-Bucket architecture. I provisioned a Google Cloud Storage (GCS) bucket and configured our FastAPI backend to generate time-boxed signed URLs. This allowed the client to upload files directly to the cloud, completely bypassing the web server. Once uploaded, a new asynchronous pipeline pulls the data and processes it in memory-safe 500-record chunks.&lt;/p&gt;

&lt;p&gt;However, when we moved to End-to-End (E2E) testing to verify this new architecture, we hit a wall.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzjqilqetrzwutei2jczx.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzjqilqetrzwutei2jczx.png" alt=" " width="800" height="270"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Root Cause Analysis:&lt;/em&gt; 404 Dashboard Errors: As the terminal output shows, the frontend UI was attempting to fetch data from /api/companies/acme/provider_walkthrough and /api/companies/acme/records. These endpoints were returning 404 Not Found because they hadn't been fully scaffolded on the FastAPI side yet.   &lt;/p&gt;

&lt;p&gt;&lt;em&gt;The 403 Upload Block:&lt;/em&gt; Concurrently, the backend strictly requires an X-User-Role HTTP header to authorize the signed URL generation. The React frontend was reading the user's admin role from the UI state but failing to attach it to the network request. Without that header, the backend defaulted to a read-only "Viewer" state and blocked the upload.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Solution: Sprint 7 and Handshake Stabilization&lt;/strong&gt;&lt;br&gt;
To properly manage this technical debt, we scoped out an Agile sprint specifically dedicated to resolving these cross-system integration bugs. For Sprint 7, we created a 15-story-point Epic: End-to-End (E2E) Integration &amp;amp; Handshake Stabilization.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft13fp9xqwdzp944oyzob.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft13fp9xqwdzp944oyzob.png" alt=" " width="800" height="212"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Here is how we divided the workload to patch the pipeline:   &lt;/p&gt;

&lt;p&gt;&lt;em&gt;Auth &amp;amp; Header Parity:&lt;/em&gt; Collin is updating the React fetch requests to explicitly pass the UI state role into the X-User-Role header, clearing the backend’s security policies and resolving the 403 upload rejection.   &lt;/p&gt;

&lt;p&gt;&lt;em&gt;Endpoint Alignment:&lt;/em&gt; I am scaffolding and wiring up the missing /records and /provider_walkthrough API routes so the client can successfully fetch and render the pipeline's output metrics without hitting the 404s we saw in the server logs.   &lt;/p&gt;

&lt;p&gt;&lt;em&gt;E2E Dynamic Load Testing:&lt;/em&gt; Once the handshake is secure, I will conduct a dynamic analysis load test, pushing a massive dataset through the application. We will monitor the backend memory consumption logs to verify that the 500-record chunking pipeline permanently resolves the previous OOM crashes.   &lt;/p&gt;

&lt;p&gt;&lt;em&gt;Graceful UI Error States:&lt;/em&gt; Instead of a generic failure message, the UI will now gracefully handle network drops and explicitly warn users if they lack the required editor permissions.   &lt;/p&gt;

&lt;p&gt;Software integration is rarely just about making the code work; it’s about making the systems communicate securely and gracefully. With this sprint board locked in, we have a clear path to a stable, production-ready upload pipeline. Check back next week for the results of our load test.&lt;/p&gt;

</description>
      <category>architecture</category>
      <category>backend</category>
      <category>performance</category>
      <category>react</category>
    </item>
    <item>
      <title>Scaling MedReachAI: Overcoming Cloud Limits with Signed URLs and Stream Processing</title>
      <dc:creator>Scott Shoemaker</dc:creator>
      <pubDate>Tue, 22 Sep 2026 14:46:15 +0000</pubDate>
      <link>https://dev.to/scott_shoemaker_8d10ccbd2/scaling-medreachai-overcoming-cloud-limits-with-signed-urls-and-stream-processing-1pa5</link>
      <guid>https://dev.to/scott_shoemaker_8d10ccbd2/scaling-medreachai-overcoming-cloud-limits-with-signed-urls-and-stream-processing-1pa5</guid>
      <description>&lt;p&gt;As my capstone partner Collin and I press forward into the final phases of MedReachAI, our focus has shifted from core feature development to production scalability and infrastructure hardening. This week, we kicked off Sprint 6, tackling one of the most critical engineering hurdles an enterprise data platform can face: safely ingesting massive, real-world datasets without crashing our server environment.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Problem:&lt;/strong&gt; Hitting the Cloud Infrastructure Wall&lt;br&gt;
MedReachAI is designed to clean, scrub, and analyze massive Healthcare Provider (HCP) datasets—often containing hundreds of thousands of records spanning over 800 MB per file. During initial end-to-end testing, we attempted to upload a full-scale national dataset and immediately ran into a brick wall of infrastructure limits:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Network Bottleneck:&lt;/strong&gt; Google Cloud Run enforces a strict 32 MiB maximum limit on incoming HTTP requests. Any file payload sent directly through our API triggers an immediate HTTP 413 error before the data even reaches our backend logic.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Memory Exhaustion (OOM Crashes):&lt;/strong&gt; Our FastAPI backend container is provisioned with 1 GiB of RAM. Attempting to parse an 800 MB CSV file synchronously into memory—while simultaneously running heavy spaCy and Presidio NLP models for PII detection and anomaly flagging—guarantees a fatal Out-of-Memory container crash.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fc336ui7klex2h78rdddw.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fc336ui7klex2h78rdddw.png" alt=" " width="799" height="335"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Solution:&lt;/strong&gt; &lt;em&gt;A Scalable Direct-to-Bucket Architecture&lt;/em&gt;&lt;br&gt;
To solve these scaling challenges, we designed and planned a complete architectural overhaul under our Large File Handling Architecture epic. Instead of routing massive file payloads through our web server, we are decoupling file transport from backend processing using a three-tiered approach:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Backend Signed URL Generation:&lt;/em&gt; We are implementing a new FastAPI endpoint utilizing the google-cloud-storage SDK. When a user initiates an upload, the backend verifies their permissions and generates a secure, time-limited v4 Signed URL pointing directly to our Google Cloud Storage (GCS) bucket.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Direct-to-Client Bucket Upload:&lt;/em&gt; Collin is refactoring the React frontend upload workflow. Instead of sending the file body to our API, the client fetches the Signed URL and executes a direct PUT request to GCS, completely bypassing the 32 MiB Cloud Run request limit and enabling progress-tracked streaming over secure channels.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Asynchronous Stream Processing:&lt;/em&gt; Once the raw file safely lands in GCS, an event trigger notifies the backend to begin processing. Rather than loading the entire dataset into RAM, scrubbing_pipeline.py is being refactored to read the file as an asynchronous data stream, processing records in memory-safe chunks to keep our container well below its 1 GiB ceiling.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsoq9fmsneaz17wtwvi34.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsoq9fmsneaz17wtwvi34.jpg" alt=" " width="800" height="343"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Looking Ahead&lt;/strong&gt;&lt;br&gt;
By shifting from direct HTTP payload transfers to a decoupled GCS storage and streaming model, MedReachAI will be fully equipped to handle enterprise-grade healthcare datasets smoothly and reliably.&lt;/p&gt;

&lt;p&gt;Stay tuned as Collin and I execute these Sprint 6 tasks and bring our capstone architecture to its production-ready finish line!&lt;/p&gt;

</description>
      <category>programming</category>
      <category>beginners</category>
      <category>fastapi</category>
      <category>react</category>
    </item>
    <item>
      <title>Sprint 5: Breaking Local Constraints and Moving MedReachAI to the Cloud</title>
      <dc:creator>Scott Shoemaker</dc:creator>
      <pubDate>Wed, 16 Sep 2026 15:16:12 +0000</pubDate>
      <link>https://dev.to/scott_shoemaker_8d10ccbd2/sprint-5-breaking-local-constraints-and-moving-medreachai-to-the-cloud-430a</link>
      <guid>https://dev.to/scott_shoemaker_8d10ccbd2/sprint-5-breaking-local-constraints-and-moving-medreachai-to-the-cloud-430a</guid>
      <description>&lt;p&gt;&lt;strong&gt;The Problem:&lt;/strong&gt; &lt;em&gt;The Local Development Bottleneck&lt;/em&gt;&lt;br&gt;
Up until this week, the MedReachAI application was strictly confined to local development environments. While this is fine for initial prototyping, it created a massive bottleneck for team collaboration. My project partner, Collin, was building the React frontend, while I was constructing the Python FastAPI backend. The immediate problem was connectivity and accessibility: the live Firebase frontend could not communicate with a backend running on localhost. Furthermore, simulating real-world latency, handling Cross-Origin Resource Sharing (CORS) security policies, and conducting end-to-end QA testing is impossible when an application isn't truly "live." We needed a scalable, cloud-hosted infrastructure that allowed independent, continuous deployment.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Solution:&lt;/strong&gt; &lt;em&gt;Serverless Containerization&lt;/em&gt;&lt;br&gt;
To solve this, Sprint 5 was entirely dedicated to cloud infrastructure migration. Instead of dealing with the overhead of heavy local container software, I opted for a lightweight, automated pipeline using Google Cloud Build and Cloud Run.&lt;/p&gt;

&lt;p&gt;Here is how we resolved the infrastructure roadblocks:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Backend Containerization:&lt;/em&gt; I wrote a custom Dockerfile to package the FastAPI application, its dependencies (including heavier libraries like spaCy and pandas), and the server execution logic into a single, reproducible image.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Automated CI/CD Pipeline:&lt;/em&gt; By linking our GitHub repository directly to Google Cloud Build, any code pushed to the backend-dev branch now automatically triggers a cloud-side compilation. This offloads the hardware requirements from my local machine directly to Google's servers.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Serverless Deployment:&lt;/em&gt; Cloud Run takes that built container and hosts it on a public HTTPS endpoint. By allocating 1 GiB of memory, we ensured the Python environment has the resources it needs without risking Out-Of-Memory (OOM) crashes on startup.&lt;/p&gt;

&lt;p&gt;_API Routing: _Finally, to bypass CORS limitations, we updated our firebase.json rewrite rules. Now, when the React frontend calls an API endpoint, Firebase securely proxies that request directly to the live Cloud Run instance.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuovkns9aay8549lyqkst.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuovkns9aay8549lyqkst.png" alt=" " width="800" height="993"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Result:&lt;/strong&gt; &lt;strong&gt;A Fully Hosted Capstone&lt;/strong&gt;&lt;br&gt;
This architecture fundamentally changed our workflow. The backend API is now officially live on the web and fully integrated. We no longer have to spin up local terminals or worry about environment discrepancies.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp7i6owtfr222oej37c6h.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp7i6owtfr222oej37c6h.png" alt=" " width="800" height="863"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Looking Ahead to QA&lt;/strong&gt;&lt;br&gt;
With the infrastructure hardened, the remainder of Sprint 5 is focused on optimizing database performance and aggressive stress testing.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1vvw8u9cmmvuveuy38yy.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1vvw8u9cmmvuveuy38yy.png" alt=" " width="800" height="232"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Our upcoming backlog includes:&lt;/p&gt;

&lt;p&gt;_Batched Writes &amp;amp; Indexing: _Optimizing bulk flag resolutions and implementing custom Firestore indexing to accelerate query speeds on the new cloud architecture.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Resilience Engineering:&lt;/em&gt; Designing asynchronous exponential backoff and retry logic to prevent data loss during network hiccups.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Stress Testing:&lt;/em&gt; Using tools like Locust and JMeter to evaluate concurrency limits, alongside rigorous vulnerability testing against malicious payloads and CSV injections.&lt;/p&gt;

&lt;p&gt;The foundation is built; now it is time to see how much stress it can handle.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Taming Messy Healthcare Data: Entity Resolution and Overcoming React's "Invisible Shield"</title>
      <dc:creator>Scott Shoemaker</dc:creator>
      <pubDate>Tue, 08 Sep 2026 12:58:14 +0000</pubDate>
      <link>https://dev.to/scott_shoemaker_8d10ccbd2/taming-messy-healthcare-data-entity-resolution-and-overcoming-reacts-invisible-shield-5hh7</link>
      <guid>https://dev.to/scott_shoemaker_8d10ccbd2/taming-messy-healthcare-data-entity-resolution-and-overcoming-reacts-invisible-shield-5hh7</guid>
      <description>&lt;p&gt;Building a compliance tool like MedReachAI means confronting one of the healthcare industry's most expensive problems: messy data. Medical registries, pharmaceutical payment logs, and clinical databases are notoriously riddled with human typos, duplicate provider entries, and missing Protected Health Information (PHI).&lt;/p&gt;

&lt;p&gt;When a compliance officer has to manually scrub a spreadsheet containing thousands of rows, the risk of human error is immense. Accidentally exporting unvalidated data can trigger severe regulatory liabilities. This week, my teammate Collin and I focused on automating this audit process by connecting our Python anomaly detection backend to our React frontend, creating a "Human-in-the-Loop" resolution dashboard.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Problem: Fragmented Data and UI Disconnects&lt;/strong&gt;&lt;br&gt;
Our primary architectural goal this week was to implement automated entity resolution. We needed the system to recognize that two separate rows with missing data (e.g., a missing email in one, a missing phone number in another) actually belong to the same provider, rather than treating them as isolated errors.&lt;/p&gt;

&lt;p&gt;While the Python backend was successfully clustering these duplicates using the provider's National Provider Identifier (NPI), we hit a significant technical roadblock on the React frontend.&lt;/p&gt;

&lt;p&gt;We built a bottom table to display flagged records, intending for the user to click a provider's name and have the top dashboard instantly snap to that provider's detailed profile. However, nothing happened when we clicked the rows. The custom table components we were using were acting as an "invisible shield," silently swallowing the native onClick events. Furthermore, the derived arrays populating the table were stripping out the original database IDs, leaving React with no way to trace the click back to the source object.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkys24fy6bt4k7c6pmrnp.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkys24fy6bt4k7c6pmrnp.png" alt=" " width="799" height="313"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Solution: Native Spans and The "Golden Record"&lt;/strong&gt;&lt;br&gt;
To resolve the state management disconnect, we had to bypass the custom UI wrapper entirely.&lt;/p&gt;

&lt;p&gt;First, we refactored the derived PII rows to explicitly preserve and pass the originalRecord object. Second, instead of attaching the click handler to the parent , we wrapped the provider's name cell in a native HTML &lt;span&gt;. By attaching the onClick event directly to this native element and utilizing e.stopPropagation(), we forced the event through the custom table wrapper.&lt;/span&gt;&lt;/p&gt;

&lt;p&gt;The moment this state connection clicked into place, the true power of our anomaly engine became visible.&lt;/p&gt;

&lt;p&gt;Now, when you click on a provider—let's say, "Riley Flores"—the dashboard doesn't just show a single row of data. Because the system anchors its logic to the NPI, it instantly pulls in all associated duplicate records and clusters them under Riley’s primary profile. We call this the "Golden Record." The compliance officer gets full context in a single view, complete with dynamically calculated null-percentages for missing emails and phone numbers, and one-click controls to safely Anonymize or Remove the bad data.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0hmvtej2aynaqskb3249.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0hmvtej2aynaqskb3249.png" alt=" " width="800" height="425"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What’s Next: Stress-Testing with CMS National Data&lt;/strong&gt;&lt;br&gt;
Up to this point, we have been validating our logic against a curated 60-row synthetic dataset. While it perfectly demonstrates our statistical anomaly detection—like catching a mathematically impossible 75,000 prescription volume—we need to prove this architecture scales.&lt;/p&gt;

&lt;p&gt;Following some excellent feedback from our professor, our next major milestone is stress-testing the pipeline. I will be ingesting the official Centers for Medicare &amp;amp; Medicaid Services (CMS) "Doctors and Clinicians" national dataset. My goal is to map this massive, real-world file to our pipeline and ensure our Python engine can process, flag, and group the data without memory timeouts before cleanly passing it to the React frontend.&lt;/p&gt;

&lt;p&gt;We have the logic locked in; now it is time to see how it handles the weight of national data.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Stepping into Beta: Explainable AI and Securing MedReachAI</title>
      <dc:creator>Scott Shoemaker</dc:creator>
      <pubDate>Tue, 01 Sep 2026 15:49:21 +0000</pubDate>
      <link>https://dev.to/scott_shoemaker_8d10ccbd2/stepping-into-beta-explainable-ai-and-securing-medreachai-218d</link>
      <guid>https://dev.to/scott_shoemaker_8d10ccbd2/stepping-into-beta-explainable-ai-and-securing-medreachai-218d</guid>
      <description>&lt;p&gt;As we officially transition from the Alpha phase of our Capstone into Beta development, the focus for MedReachAI is shifting from core functionality to real-world usability and enterprise-grade security.&lt;/p&gt;

&lt;p&gt;Over the last few weeks, Collin and I successfully built a backend data intelligence pipeline capable of processing healthcare provider data, identifying statistical anomalies using machine learning, and verifying regulatory compliance against federal databases. But as we map out our next two-week sprint, we are preparing to tackle two massive hurdles: making our AI explainable to human stakeholders, and locking down the system architecture.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Problem: The "Black Box" of AI&lt;/strong&gt;&lt;br&gt;
In our latest milestone review, we received fantastic feedback: the underlying architecture works, and the pipeline correctly flags bad data. However, there was a critical piece missing. Stakeholders don't just want to know that a record was flagged; they need to know why.&lt;/p&gt;

&lt;p&gt;When an unsupervised machine learning model (like our IsolationForest) outputs a -1 to indicate an anomaly, it creates a "black box" scenario. For a healthcare data steward tasked with managing critical provider directories, seeing a mathematical variance score isn't actionable. If users don’t understand the AI’s decision-making process, they won't trust the platform.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Solution: "Day in the Life" Walkthrough &amp;amp; Explainable AI&lt;/strong&gt;&lt;br&gt;
To solve this ambiguity, we are dedicating the first week of this sprint to building an interactive visualization that demystifies the AI pipeline. We plan to engineer a "Plain-Language Anomaly Explanation Generator." Instead of just returning a variance score, our goal is for the backend to dynamically translate the mathematical anomaly into human-readable context (e.g., "Flagged because billing volume is 5x the Cardiology average").&lt;/p&gt;

&lt;p&gt;To build and test this feature, I will be writing a Python script to generate a highly curated synthetic dataset. We intend to intentionally inject four distinct "bad data" personas into this batch to test every layer of our pipeline:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The Statistical Anomaly&lt;/em&gt;: A provider with massive prescription volumes to test the machine learning model and trigger the future plain-text generator.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The NPI Failure&lt;/em&gt;: A deactivated National Provider Identifier to ensure our CMS validation flag catches it.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The Financial Conflict&lt;/em&gt;: A provider exceeding the $100 Sunshine Act limit to trigger our local SQLite compliance database.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The Duplicate&lt;/em&gt;: An identical record to trigger our Pandas soft-delete flag.&lt;/p&gt;

&lt;p&gt;Once this dataset is generating properly, Collin will build out a "Record Detail" UI in React. This will allow us to step through these exact records and visually demonstrate the pipeline's logic to our stakeholders in real-time.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F93h6ac4nv7wk96jbzx1a.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F93h6ac4nv7wk96jbzx1a.png" alt=" " width="800" height="331"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Problem: Securing Healthcare Infrastructure&lt;/strong&gt;&lt;br&gt;
The second major hurdle for our Beta release is security. Up to this point, our local environment allowed open communication between our React frontend and FastAPI backend. In a production healthcare environment, leaving endpoints unauthenticated is an absolute non-starter. We need strict access controls to ensure that only authorized personnel can view or manipulate data, and that data from one organization cannot bleed into another.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Solution: JWT Middleware and RBAC&lt;/strong&gt;&lt;br&gt;
For the second half of this sprint, my focus will shift entirely to Phase C architecture: Security and Authentication.&lt;/p&gt;

&lt;p&gt;I plan to implement a robust Firebase Auth JSON Web Token (JWT) Middleware layer on the FastAPI backend. Every single API request will require a verified token before the server even looks at the data. Once the perimeter is secure, we will roll out Role-Based Access Control (RBAC).&lt;/p&gt;

&lt;p&gt;Admins will have full access to team management and data scrubbing.&lt;/p&gt;

&lt;p&gt;Editors can review and flag data.&lt;/p&gt;

&lt;p&gt;Viewers will be strictly read-only.&lt;/p&gt;

&lt;p&gt;Coupled with upcoming multi-tenant Firestore security rules, this will ensure our platform is isolated, secure, and ready for a true Beta deployment.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fl46tugei9n191k3ix6g2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fl46tugei9n191k3ix6g2.png" alt=" " width="800" height="302"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Moving Forward&lt;/strong&gt;&lt;br&gt;
This two-week hybrid sprint is designed to perfectly bridge the gap between demonstrating immediate value to our stakeholders and building the invisible, architectural security foundation required for a production-ready application. I am looking forward to seeing this "Day in the Life" walkthrough come to life and reporting back on how the plain-language generator handles the synthetic data next week.&lt;/p&gt;

</description>
      <category>healthtech</category>
      <category>ai</category>
      <category>explainableai</category>
      <category>fastapi</category>
    </item>
    <item>
      <title>MedReachAI: Bringing Data Intelligence to Healthcare Provider Management</title>
      <dc:creator>Scott Shoemaker</dc:creator>
      <pubDate>Thu, 27 Aug 2026 21:02:52 +0000</pubDate>
      <link>https://dev.to/scott_shoemaker_8d10ccbd2/medreachai-bringing-data-intelligence-to-healthcare-provider-management-4lmb</link>
      <guid>https://dev.to/scott_shoemaker_8d10ccbd2/medreachai-bringing-data-intelligence-to-healthcare-provider-management-4lmb</guid>
      <description>&lt;p&gt;&lt;strong&gt;The Messy Reality of Healthcare Data &lt;em&gt;(The Problem)&lt;/em&gt;&lt;/strong&gt;&lt;br&gt;
When people hear the phrase "healthcare data," they immediately think of patient records—blood pressure, clinical notes, and heart rates. But there is an entirely different side to the industry that is just as critical, and often just as messy: Provider Data.&lt;/p&gt;

&lt;p&gt;Healthcare organizations, compliance officers, and medical sales teams rely on massive datasets of physician credentials, practice locations, and financial relationships. However, this data is notoriously unstructured. It suffers from a few major problems:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Corruption &amp;amp; Duplication:&lt;/strong&gt; Provider databases are often cobbled together through web-scraping or manual entry. This leads to duplicate profiles, mismatched specialties, and corrupted National Provider Identifiers (NPIs).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hidden Security Risks:&lt;/strong&gt; Free-text fields (like notes or contact arrays) frequently harbor Protected Health Information (PHI) or Personally Identifiable Information (PII), creating massive compliance liabilities.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Financial Blind Spots:&lt;/strong&gt; Without a clean, verified NPI, it is impossible to track a physician's financial relationships under the federal Sunshine Act.&lt;/p&gt;

&lt;p&gt;The industry needs a way to automatically ingest, sanitize, and secure this provider data before it ever hits a production database.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvemtk5ia38vay3jh1sw4.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvemtk5ia38vay3jh1sw4.jpg" alt=" " width="799" height="436"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;(Caption: Raw healthcare provider datasets are often riddled with structural errors and hidden compliance risks.)&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Enter MedReachAI &lt;em&gt;(The Solution)&lt;/em&gt;&lt;/strong&gt;&lt;br&gt;
Over the past few months, my capstone partner, Collin, and I have been building MedReachAI: an automated, full-stack data intelligence platform designed specifically to sanitize and enrich healthcare provider profiles.&lt;/p&gt;

&lt;p&gt;Rather than relying on humans to manually comb through spreadsheets, we engineered a backend pipeline using Python, FastAPI, and machine learning to do the heavy lifting.&lt;/p&gt;

&lt;p&gt;Here is a look at the core features driving the MedReachAI backend:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;em&gt;Unsupervised Anomaly Detection&lt;/em&gt;&lt;br&gt;
To catch corrupted data, I engineered a machine learning model using Scikit-Learn’s IsolationForest. Because provider data changes dynamically, we don't rely on a static training set. Instead, the algorithm evaluates each new upload batch in real-time, calculates the mathematical distance between provider profiles, and automatically isolates the top 5% most irregular records (such as geographic mismatches or structural scraping errors).&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;em&gt;Automated PII &amp;amp; PHI Redaction&lt;/em&gt;&lt;br&gt;
To handle security, I integrated the Microsoft Presidio AnalyzerEngine. I built custom pattern recognizers that scan unstructured text to detect Medical Record Numbers (MRNs) alongside standard identifiers like SSNs and emails. Once detected, our downstream Anonymizer service actively scrubs the data, replacing the sensitive text with explicit tags like .&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpc8w384440i2p0y1e6n0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpc8w384440i2p0y1e6n0.png" alt=" " width="800" height="649"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;(Caption: The MedReachAI backend utilizes custom Python services to flag statistical anomalies and redact sensitive information in real-time.)&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;em&gt;Non-Destructive Data Deduplication&lt;/em&gt;&lt;br&gt;
To ensure our analytics remain accurate, I built a custom Pandas algorithm to identify exact and partial duplicate provider profiles. Instead of deleting these records and risking data loss, the system performs a "soft delete" by appending a boolean flag to the database row, allowing the data to be recovered if needed.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;em&gt;The Human-in-the-Loop Dashboard&lt;/em&gt;&lt;br&gt;
While the backend handles the automation, Collin built a brilliant React-based UI that gives the user final oversight. The MedReachAI dashboard features a suite of review tabs that allow end-users to manually accept or reject the anomaly, duplication, and PII flags generated by the backend.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvbtmkxv0ew9w77zjx3ui.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvbtmkxv0ew9w77zjx3ui.png" alt=" " width="799" height="611"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;(Caption: The MedReachAI dashboard provides end-users with real-time visibility into the health and cleanliness of their provider datasets.)&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What’s Next?&lt;/strong&gt;&lt;br&gt;
With the core data-cleaning pipeline fully operational, Phase C of our development will focus on external enrichment. I will be building asynchronous clients to connect our clean local records directly to the official CMS NPI Registry and the CMS Open Payments database, turning sparse, sanitized records into rich, compliant provider profiles.&lt;/p&gt;

&lt;p&gt;Building MedReachAI has been an incredible exercise in bridging the gap between raw data engineering and user-facing application design. Stay tuned as we wrap up our final features!&lt;/p&gt;

</description>
    </item>
    <item>
      <title>MedReach-AI: Kicking Off Sprint 3 and the "Brain" of Phase B</title>
      <dc:creator>Scott Shoemaker</dc:creator>
      <pubDate>Tue, 18 Aug 2026 11:01:22 +0000</pubDate>
      <link>https://dev.to/scott_shoemaker_8d10ccbd2/medreach-ai-kicking-off-sprint-3-and-the-brain-of-phase-b-23c5</link>
      <guid>https://dev.to/scott_shoemaker_8d10ccbd2/medreach-ai-kicking-off-sprint-3-and-the-brain-of-phase-b-23c5</guid>
      <description>&lt;p&gt;Hey everyone, Scott here. Collin and I are back with the latest engineering update for MedReach-AI.&lt;/p&gt;

&lt;p&gt;We just officially closed out Phase A of development. Over the last two weeks, we built a rock-solid foundation for our file ingestion pipelines, set up a heuristic engine that automatically maps healthcare columns, and locked down our export endpoints with strict HIPAA-compliant filters.&lt;/p&gt;

&lt;p&gt;With the plumbing working flawlessly, we are now shifting into Phase B: Data Intelligence. This is a massive two-week sprint where the workload essentially doubles, and we transition from simply moving data around to actively analyzing and sanitizing it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fje4hebdyne2bker39fon.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fje4hebdyne2bker39fon.png" alt=" " width="800" height="422"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Sprint 3 Goal&lt;/strong&gt;&lt;br&gt;
To put it simply: we are building the brain. Our goal for Sprint 3 is to take raw, uploaded healthcare datasets and run them through automated PII/PHI scrubbing, real-time government registry validations, and machine learning anomaly detection before the user ever sees the final dashboard.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fseodc9u38jh6pyyf8oih.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fseodc9u38jh6pyyf8oih.png" alt=" " width="799" height="329"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fez556tgrvvqvbrtv9n29.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fez556tgrvvqvbrtv9n29.png" alt=" " width="799" height="139"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How We Are Tackling Phase B&lt;/strong&gt;&lt;br&gt;
Looking at our Jira board, we have broken this phase down into a few core pillars:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Machine Learning (The AI in MedReach-AI)&lt;/strong&gt;: This is the part I am most excited to dive into. We don't have months to train a massive neural network, so we are utilizing unsupervised machine learning via Scikit-Learn. By deploying IsolationForest and DBSCAN clustering algorithms, our backend will be able to mathematically evaluate a dataset on the fly and flag statistical outliers or corrupted records without relying on rigid, hard-coded rules.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Security &amp;amp; Compliance&lt;/strong&gt;: We are integrating the Microsoft Presidio engine to actively detect and anonymize protected health information (PHI). We are also building ingestion engines for CMS Open Payments data to flag financial conflicts of interest (Sunshine Act compliance).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;External Validation&lt;/strong&gt;: We are building an asynchronous client to query the official CMS NPI Registry. This will validate provider credentials in real-time and enrich our database with verified classification statuses.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Dashboards &amp;amp; Deduplication&lt;/strong&gt;: While I handle the machine learning and deduplication algorithms on the backend, Collin is building out the interactive React tabs. These interfaces will allow compliance officers to visually review our AI’s flags, merge duplicate records, and download legally sound PDF audit trails.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What’s Next?&lt;/strong&gt;&lt;br&gt;
It is a heavy workload, but our Phase A architecture is clean, and we are ready to execute. For the first week of this sprint, I am focusing entirely on getting that Scikit-Learn Isolation Forest up and running to start flagging our test data.&lt;/p&gt;

&lt;p&gt;We will be back next week with a mid-sprint update to share how the anomaly models are performing!&lt;/p&gt;

</description>
    </item>
    <item>
      <title>MedReach-AI: Sprint 2 Progress &amp; Milestone 2.1 Update</title>
      <dc:creator>Scott Shoemaker</dc:creator>
      <pubDate>Fri, 14 Aug 2026 23:04:45 +0000</pubDate>
      <link>https://dev.to/scott_shoemaker_8d10ccbd2/medreach-ai-sprint-2-progress-milestone-21-update-621</link>
      <guid>https://dev.to/scott_shoemaker_8d10ccbd2/medreach-ai-sprint-2-progress-milestone-21-update-621</guid>
      <description>&lt;p&gt;That makes perfect sense. Since Phase A is a full two-week sprint (August 10th to August 24th), today (August 14th) marks the exact midpoint.&lt;/p&gt;

&lt;p&gt;Framing this as a midway check-in adds a lot of professionalism. It shows you aren't just scrambling week-to-week, but executing a well-planned, two-week architecture. Week 1 was about safely catching the data; Week 2 is about aggressively processing it.&lt;/p&gt;

&lt;p&gt;Here is the updated blog post reflecting the two-week Phase A narrative.&lt;/p&gt;

&lt;p&gt;MedReach-AI: Phase A (Mid-Sprint) Progress &amp;amp; Milestone 2.1 Update&lt;br&gt;
Welcome back to the MedReach-AI development log. I’m Scott Shoemaker, alongside my co-developer Collin. As we push through our software engineering capstone, we are excited to share our progress for the 2.1 Project Milestone.&lt;/p&gt;

&lt;p&gt;We are currently executing Phase A: Data Management, which is scoped as a comprehensive two-week sprint running from August 10th to August 24th. Today marks the midpoint of this sprint. For Week 1, our engineering focus was establishing the core file upload architecture on both the frontend interface and the backend API.&lt;/p&gt;

&lt;p&gt;Here is a breakdown of what we accomplished in the first half of Phase A, the roadblocks we cleared, and our roadmap for Week 2.&lt;/p&gt;

&lt;p&gt;Week 1 Backend: Securing the Data Pipeline&lt;br&gt;
On the backend, my primary responsibility for this first week was ensuring our FastAPI server could safely catch and store incoming data streams before we begin processing them. I completed two major tickets:&lt;/p&gt;

&lt;p&gt;MA-17: Firebase Storage Upload Endpoint (Completed): I engineered a tenant-scoped file upload endpoint. When a file is uploaded, the API successfully routes it to our local Firebase Storage emulator. It returns a 201 Created status with a generated UUID and the definitive storage path, guaranteeing strict data isolation between clinics.&lt;/p&gt;

&lt;p&gt;MA-22: Malformed Data &amp;amp; Schema Error Wrapper (Completed): To ensure bad data doesn't crash the server during Week 2's heavy processing, I implemented a validation wrapper using Python's csv module. If a clinic uploads a corrupted file, the server intercepts the unparseable delimiters or decoding errors and returns a structured 400 Bad Request. This payload identifies the exact line number of the error, keeping the application stable.&lt;/p&gt;

&lt;p&gt;Week 1 Frontend: Building the Interface&lt;br&gt;
On the frontend, Collin spent this week translating our Phase A wireframes into functional, interactive React components:&lt;/p&gt;

&lt;p&gt;MA-16: UI Drag-and-Drop File Upload Component (Completed): Users now have a clean, intuitive dropzone to pull their CSV files into the application.&lt;/p&gt;

&lt;p&gt;MA-20: UI Column Mapping Confirmation &amp;amp; Override Table (Completed): Collin built out the data table where users can visually review their uploaded data and manually override data types before final submission. The state management is wired up locally and prepped for backend integration.&lt;/p&gt;

&lt;p&gt;MA-15: UI Export Screen &amp;amp; File Download Controls (In Progress): The core layout for the data export screen is established. Finishing the button states and download triggers will carry over into early next week.&lt;/p&gt;

&lt;p&gt;Roadblocks &amp;amp; Technical Challenges&lt;br&gt;
No sprint is without its hurdles. On the backend, setting up the storage emulator for MA-17 resulted in a massive port conflict. Localhost ports 8080 and 8000 were clashing with background processes, crashing our startup sequences. I resolved this by remapping the Firestore emulator to port 8081 and updating our configurations, successfully clearing the network blockage.&lt;/p&gt;

&lt;p&gt;On the frontend, ensuring the local React state perfectly matches the expected backend JSON payloads took some extra coordination. Because of this, MA-15 required a bit more time and will bridge the gap into Week 2 to ensure the data formatting is perfectly aligned.&lt;/p&gt;

&lt;p&gt;The Roadmap: Week 2 of Phase A&lt;br&gt;
With the file ingestion infrastructure stabilized in Week 1, the second half of this sprint is entirely dedicated to active data processing and endpoint integration.&lt;/p&gt;

&lt;p&gt;Scott's Backend Tasks (Week 2):&lt;/p&gt;

&lt;p&gt;MA-21: Pandas Chunked CSV Ingestion Engine: Standard of Completion: The engine must successfully process a 100,000-row CSV file without exceeding 512 megabytes of server RAM. I will utilize the Pandas chunksize parameter and tracemalloc to strictly monitor and prove our memory footprint.&lt;/p&gt;

&lt;p&gt;MA-19: Heuristic Column Type Detection: Standard of Completion: The backend must automatically evaluate the sample data and infer column types (like emails and phone numbers) using regex, returning a mapped schema dictionary.&lt;/p&gt;

&lt;p&gt;MA-18: Clean CRM-Compatible CSV Export Pipeline: Standard of Completion: The API successfully returns a sanitized dataset formatted for standard CRM ingestion.&lt;/p&gt;

&lt;p&gt;Collin's Frontend Tasks (Week 2):&lt;/p&gt;

&lt;p&gt;Complete MA-15: Standard of Completion: A fully functional download interface that triggers the browser's native file save dialog with the scrubbed data.&lt;/p&gt;

&lt;p&gt;Endpoint Integration: Wiring the completed Phase A frontend React components directly to the active FastAPI endpoints to establish end-to-end functionality.&lt;/p&gt;

&lt;p&gt;We are in a great position at the midway point of Phase A. Our repositories are stable, our data validation is bulletproof, and our interface is coming to life. Check back at the end of the sprint as we fire up the data-scrubbing engine!&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Sprint 1: Standing Up the MedReach-AI Backend (and Taming Emulators)</title>
      <dc:creator>Scott Shoemaker</dc:creator>
      <pubDate>Fri, 07 Aug 2026 01:27:21 +0000</pubDate>
      <link>https://dev.to/scott_shoemaker_8d10ccbd2/sprint-1-standing-up-the-medreach-ai-backend-and-taming-emulators-589c</link>
      <guid>https://dev.to/scott_shoemaker_8d10ccbd2/sprint-1-standing-up-the-medreach-ai-backend-and-taming-emulators-589c</guid>
      <description>&lt;p&gt;Kicking off a brand-new project is always a fun mix of grand architecture planning and immediately running into weird setup bugs. This week, we officially started the MedReach-AI capstone project, and my main focus was getting our Python and FastAPI backend repository scaffolded and ready to go.&lt;/p&gt;

&lt;p&gt;Overall, the basic server setup was a breeze, but hooking up our local dev tools and establishing the Git workflow threw a couple of curveballs my way.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Problem: Firebase Emulators vs. Strict Security&lt;/strong&gt;&lt;br&gt;
I set up the Firebase Emulator Suite for our Firestore database so we can develop locally without worrying about cloud costs. Usually, when you connect the Firebase Admin SDK to a local emulator, you just pass a dummy credential dictionary with a fake private key (like "private_key": "mock-key"). But when I spun up the Uvicorn server to test it out, the app immediately crashed.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhp7elddq4gf5qixgxv4c.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhp7elddq4gf5qixgxv4c.png" alt=" " width="800" height="66"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;It turns out the underlying Python cryptography package recently got a security upgrade. It now strictly validates the math behind RSA keys before the app is even allowed to start. Since my "mock-key" obviously wasn't a valid RSA structure, it threw a ValueError: Unable to load PEM file and stopped the server dead in its tracks.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Fix: Demo Projects to the Rescue&lt;/strong&gt;&lt;br&gt;
Instead of wasting time trying to generate a mathematically valid fake RSA key—which kind of defeats the purpose of a quick local setup—I realized I could bypass the certificate requirement entirely.&lt;/p&gt;

&lt;p&gt;Firebase actually has a built-in way to handle this. By stripping out the dummy credential and initializing the app with a "demo" project ID, the SDK knows it's strictly in an offline testing environment.&lt;/p&gt;

&lt;p&gt;Python&lt;br&gt;
Tell the SDK to connect to the local emulator&lt;br&gt;
os.environ["FIRESTORE_EMULATOR_HOST"] = "127.0.0.1:8080"&lt;br&gt;
Initialize using a demo project ID instead of a certificate&lt;br&gt;
if not firebase_admin._apps:&lt;br&gt;
    firebase_admin.initialize_app(options={&lt;br&gt;
        'projectId': 'demo-medreachai'&lt;br&gt;
    })&lt;/p&gt;

&lt;p&gt;That completely fixed the crash. My /api/health endpoint was instantly able to query the local Firestore database.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Problem: Git Remote Collisions&lt;/strong&gt;&lt;br&gt;
Next up, I needed to push this clean, working baseline (complete with Black and Flake8 formatting) up to our backend-dev branch on GitHub. Simple enough, right? Nope. Git immediately rejected the push.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7jn4zs1zt6sio5gdr4eu.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7jn4zs1zt6sio5gdr4eu.png" alt=" " width="798" height="57"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Because the repository was originally created on GitHub's website before I scaffolded the local files, it had an auto-generated template commit (like a default README) that my computer didn't have. Git’s safety features kicked in to stop me from accidentally wiping out that remote work.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Fix: A Good Old-Fashioned Force Push&lt;/strong&gt;&lt;br&gt;
Usually, running git pull --allow-unrelated-histories is the move here to merge the two timelines. But since my local environment was specifically built to be the absolute "source of truth" for the backend, I didn't want to merge in random GitHub template files.&lt;/p&gt;

&lt;p&gt;I went with a force push to overwrite the remote branch and establish my local folder as the true baseline:&lt;/p&gt;

&lt;p&gt;git push -u origin backend-dev --force&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Wrapping Up Sprint 1&lt;/strong&gt;&lt;br&gt;
With Git sorted out and Firebase playing nice with the emulator, the MedReach-AI backend is officially up and running. Our Uvicorn server is live, the interactive Swagger UI is generating our docs automatically, and our linters are keeping the code clean.&lt;/p&gt;

&lt;p&gt;The foundation is set, and next week we finally get to dive into feature development!&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5ygv92vellajuxp9hidi.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5ygv92vellajuxp9hidi.png" alt=" " width="800" height="424"&gt;&lt;/a&gt;&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Sprint One: Slicing the Monolith and Standing Up the Phase A Backend</title>
      <dc:creator>Scott Shoemaker</dc:creator>
      <pubDate>Sat, 25 Jul 2026 08:32:44 +0000</pubDate>
      <link>https://dev.to/scott_shoemaker_8d10ccbd2/sprint-one-slicing-the-monolith-and-standing-up-the-phase-a-backend-gcg</link>
      <guid>https://dev.to/scott_shoemaker_8d10ccbd2/sprint-one-slicing-the-monolith-and-standing-up-the-phase-a-backend-gcg</guid>
      <description>&lt;p&gt;Welcome back to the dev log. Last week, Collin and I officially pivoted MedReach AI into a "Data Intelligence First" platform, locking in our enterprise architecture and securing our multi-tenant data models. This week, the planning phase ended. Sprint 1 officially kicked off, marking our transition into active software engineering.&lt;/p&gt;

&lt;p&gt;However, as we moved to pull our Phase A (Data Management MVP) tasks from the backlog, we immediately encountered a classic software development trap.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Problem: The "8-Point Ticket" Trap&lt;/strong&gt;&lt;br&gt;
As the Scrum Master for this project, my goal is to protect our team’s velocity and prevent burnout. While planning Sprint 1, I noticed several of our backend architectural tickets—such as integrating Microsoft Presidio for PII scrubbing and building our async CMS NPI Registry validation pipeline—were estimated at 8 Story Points.&lt;/p&gt;

&lt;p&gt;In Agile development, especially for a two-man capstone team balancing outside lives and careers, an 8-point ticket is functionally a monolith. It represents too much ambiguity. If a single ticket takes 20+ hours to complete, it risks sitting in the "In Progress" column for the entire sprint. If I hit a bug on day two, my velocity flatlines, and worse, I become a bottleneck for my frontend partner who is waiting on that API contract.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnohlzbeh0gkyriq05rc9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnohlzbeh0gkyriq05rc9.png" alt=" " width="800" height="560"&gt;&lt;/a&gt;&lt;em&gt;Caption: Breaking down our Phase A backlog into highly granular, sprint-ready deliverables.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Solution: The 4-Point Hard Cap&lt;/strong&gt;&lt;br&gt;
To solve this, I enforced a strict Agile capacity rule for our board: No single task can exceed 4 Story Points.&lt;/p&gt;

&lt;p&gt;We mapped our estimation scale directly to our weekly capacity, where 1 point equals roughly 4 to 5 hours of engineering work. By capping tickets at 4 points, we ensure that the absolute maximum size of any task represents one week of part-time capacity (approx. 20 hours).&lt;/p&gt;

&lt;p&gt;I spent the first part of this week slicing our backend monoliths into testable, bite-sized components. For example, instead of one massive ticket for "NPI Validation," we now have separate 2- and 3-point tickets for the base REST connection, the async batching queue, and the local caching layer. This granularity guarantees that even if we hit a roadblock, we can still merge smaller victories and keep the burndown chart moving.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Execution: Architecting the Phase A Backend&lt;/strong&gt;&lt;br&gt;
With our sprint scope cleanly defined, I swapped my Scrum Master hat for my Backend Architect hat. My primary technical focus this week was our most critical MVP blocker: &lt;strong&gt;Security and Ingestion&lt;/strong&gt;.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;JWT Middleware &amp;amp; Multi-Tenant Isolation&lt;/strong&gt;: Before parsing a single CSV, we had to ensure our data boundaries were airtight. I mapped out and configured the Firebase Authentication middleware to intercept API requests and validate custom Role-Based Access Control (RBAC) claims. This ensures that an Editor in Company A can never access, query, or mutate data belonging to Company B at the database level.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Chunked Ingestion Pipeline&lt;/strong&gt;: I also began scaffolding the pandas chunked CSV reader. Because enterprise healthcare datasets can be massive, this async ingestion engine is designed to parse large files without maxing out our server's memory, passing the data smoothly to our heuristic column detection logic.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr77ig6y1f5f1qe0blerw.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr77ig6y1f5f1qe0blerw.png" alt=" " width="" height=""&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Caption: The MedReach AI Multi-Tenant Isolation &amp;amp; Role-Based Access Control architecture, defining strict database-level boundaries and the permission matrix for our Phase A MVP.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What’s Next?&lt;/strong&gt;&lt;br&gt;
By slicing our tasks down and securing the gateway, we have created a safe, isolated sandbox to begin heavy data manipulation. Next week, I will be deep-diving into the Microsoft Presidio integration to automate our PII/PHI redaction engine.&lt;/p&gt;

&lt;p&gt;Until next time, keep your sprints short and your endpoints secure!&lt;/p&gt;

</description>
      <category>ai</category>
      <category>career</category>
      <category>agile</category>
      <category>softwareengineering</category>
    </item>
    <item>
      <title>Sprint One: Why We Pivoted MedReach AI to "Data Intelligence First" and Killed the Paid API</title>
      <dc:creator>Scott Shoemaker</dc:creator>
      <pubDate>Sun, 19 Jul 2026 15:19:19 +0000</pubDate>
      <link>https://dev.to/scott_shoemaker_8d10ccbd2/sprint-one-why-we-pivoted-medreach-ai-to-data-intelligence-first-and-killed-the-paid-api-44b0</link>
      <guid>https://dev.to/scott_shoemaker_8d10ccbd2/sprint-one-why-we-pivoted-medreach-ai-to-data-intelligence-first-and-killed-the-paid-api-44b0</guid>
      <description>&lt;p&gt;Welcome back to the dev log! If you read my Sprint Zero post, you know my capstone partner, Collin, and I set out to build a full-stack AI platform using an asynchronous, horizontally split workflow. As we transitioned from initial scoping into Month 2 of our build, our concept matured into MedReach AI—a specialized B2B healthcare SaaS platform designed for mid-size pharmaceutical and medical device companies.&lt;/p&gt;

&lt;p&gt;However, as we dove into our software architecture review this week, we hit a massive strategic crossroads that forced us to completely refactor our engineering roadmap and Jira backlog.  &lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Problem: The Insecure "Marketing Automation" Trap&lt;/strong&gt;&lt;br&gt;
Originally, we visualized MedReach AI as an AI-driven marketing campaign generator that happened to clean Healthcare Professional (HCP) databases. But after evaluating industry pain points, we realized we were solving the wrong problem first.&lt;br&gt;&lt;br&gt;
In the pharmaceutical and medical device sectors, marketing teams don't struggle to generate email copy; they struggle with dirty, unvalidated, and legally hazardous data. Up to 30% of standard HCP databases are riddled with outdated contacts, retired physicians, and unvalidated National Provider Identifier (NPI) numbers. Running automated campaigns on bad data doesn't just waste budget—it creates massive legal exposure under federal regulations like the Physician Payments Sunshine Act and FDA off-label promotion guidelines.&lt;br&gt;&lt;br&gt;
Furthermore, enterprise platforms that actually solve this (like Veeva or IQVIA) cost upwards of $500,000 annually, pricing out two- to three-person marketing teams.  From an engineering perspective, treating data cleaning as a simple preprocessing step for a marketing tool was a recipe for disaster. It meant we were treating security, multi-tenant data isolation, and role-based access control (RBAC) as secondary features. In healthcare SaaS, if you bolt on security at the end, your platform is inherently broken.  &lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fztj3st1ei5eijjy6u8xd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fztj3st1ei5eijjy6u8xd.png" alt=" " width="799" height="538"&gt;&lt;/a&gt;&lt;em&gt;Caption: Restructuring our Jira backlog into three distinct functional milestones: Data Management, Data Intelligence, and Visualization &amp;amp; Export.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Solution: A "Data Intelligence First" Architecture&lt;/strong&gt;&lt;br&gt;
To fix this, Collin and I leaned heavily into our Agile framework and initiated a comprehensive architecture pivot. We repositioned MedReach AI entirely around Data Intelligence, treating campaign generation simply as one of several downstream outputs of a clean, standardized dataset.&lt;br&gt;&lt;br&gt;
To execute this without losing sprint velocity, we made three critical architectural decisions this week:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Shifting RBAC and Security to the MVP Blocker&lt;/strong&gt;&lt;br&gt;
We completely reorganized our 54-task Jira backlog into three new phased milestones: &lt;strong&gt;Phase A (Data Management MVP), Phase B (Data Intelligence Alpha), and Phase C (Visualization &amp;amp; Export Beta)&lt;/strong&gt;.&lt;br&gt;
Crucially, we pulled our Firebase Authentication JWT middleware, Firestore database-level multi-tenant security rules, and full-stack RBAC (Admin, Editor, Viewer roles) out of future sprints and pushed them directly into Phase A. Why? Because no file upload or cleaning pipeline should ever be tested without strict tenant isolation already enforcing database boundaries. &lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Killing the Paid API (Why Local Open-Source Wins)&lt;/strong&gt;&lt;br&gt;
In an era where every startup simply wraps the OpenAI or Anthropic API, we made a deliberate engineering choice: &lt;strong&gt;we are using zero paid external LLM APIs.&lt;/strong&gt;&lt;br&gt;
Instead, our FastAPI backend orchestrates self-hosted &lt;strong&gt;Meta Llama 3 8B&lt;/strong&gt; and &lt;strong&gt;Mistral 7B&lt;/strong&gt; models running locally on an on-premise Ollama runtime. In healthcare, clients are legally and commercially terrified of transmitting Protected Health Information (PHI) or proprietary HCP lists to third-party AI vendors. By keeping 100% of our AI inference inside our own platform infrastructure, we eliminate per-call API costs at scale and turn data privacy into our primary B2B selling point.  &lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzwrnhx0owni2pz2t3dp7.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzwrnhx0owni2pz2t3dp7.png" alt=" " width="800" height="449"&gt;&lt;/a&gt;&lt;em&gt;Caption: MedReach AI's system architecture, highlighting our self-hosted AI runtime and isolated Firestore multi-tenant data layer.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Layering the Tech Stack for Asynchronous Speed&lt;/strong&gt;&lt;br&gt;
By stabilizing our Phase A core this week, Collin and I can safely maintain our horizontal split. While Collin builds out the React 18 / Tailwind CSS drag-and-drop upload zones and natural language query interfaces on the frontend, I am building the backend Python pipelines. I'm currently integrating Microsoft Presidio for automated PII/PHI detection, scikit-learn for Isolation Forest anomaly detection, and asynchronous HTTP clients to validate provider records against the public CMS NPI Registry in real time.  &lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What This Means for My Journey&lt;/strong&gt;&lt;br&gt;
As someone aspiring to step into an AI Solutions Program Director or Product Manager role, this week was a masterclass in product framing and technical risk management. It reaffirmed a core lesson: good software architecture isn't just about writing clean code; it's about aligning your database schemas and sprint backlogs with the actual commercial and regulatory realities of your industry.&lt;br&gt;
With our roadmap refactored, our Jira board aligned with our design documentation, and our security gates locked in place, we are ready to build the core cleaning engine.&lt;br&gt;
Next stop: getting our pandas chunked CSV ingestion pipeline and Presidio anonymizers communicating smoothly. Stay tuned for Sprint 2!&lt;/p&gt;

</description>
      <category>ai</category>
      <category>healthcare</category>
      <category>fastapi</category>
      <category>career</category>
    </item>
  </channel>
</rss>
