DEV Community

Scott Shoemaker
Scott Shoemaker

Posted on

Scaling MedReachAI: Overcoming Cloud Limits with Signed URLs and Stream Processing

As my capstone partner Collin and I press forward into the final phases of MedReachAI, our focus has shifted from core feature development to production scalability and infrastructure hardening. This week, we kicked off Sprint 6, tackling one of the most critical engineering hurdles an enterprise data platform can face: safely ingesting massive, real-world datasets without crashing our server environment.

The Problem: Hitting the Cloud Infrastructure Wall
MedReachAI is designed to clean, scrub, and analyze massive Healthcare Provider (HCP) datasets—often containing hundreds of thousands of records spanning over 800 MB per file. During initial end-to-end testing, we attempted to upload a full-scale national dataset and immediately ran into a brick wall of infrastructure limits:

The Network Bottleneck: Google Cloud Run enforces a strict 32 MiB maximum limit on incoming HTTP requests. Any file payload sent directly through our API triggers an immediate HTTP 413 error before the data even reaches our backend logic.

Memory Exhaustion (OOM Crashes): Our FastAPI backend container is provisioned with 1 GiB of RAM. Attempting to parse an 800 MB CSV file synchronously into memory—while simultaneously running heavy spaCy and Presidio NLP models for PII detection and anomaly flagging—guarantees a fatal Out-of-Memory container crash.

The Solution: A Scalable Direct-to-Bucket Architecture
To solve these scaling challenges, we designed and planned a complete architectural overhaul under our Large File Handling Architecture epic. Instead of routing massive file payloads through our web server, we are decoupling file transport from backend processing using a three-tiered approach:

Backend Signed URL Generation: We are implementing a new FastAPI endpoint utilizing the google-cloud-storage SDK. When a user initiates an upload, the backend verifies their permissions and generates a secure, time-limited v4 Signed URL pointing directly to our Google Cloud Storage (GCS) bucket.

Direct-to-Client Bucket Upload: Collin is refactoring the React frontend upload workflow. Instead of sending the file body to our API, the client fetches the Signed URL and executes a direct PUT request to GCS, completely bypassing the 32 MiB Cloud Run request limit and enabling progress-tracked streaming over secure channels.

Asynchronous Stream Processing: Once the raw file safely lands in GCS, an event trigger notifies the backend to begin processing. Rather than loading the entire dataset into RAM, scrubbing_pipeline.py is being refactored to read the file as an asynchronous data stream, processing records in memory-safe chunks to keep our container well below its 1 GiB ceiling.

Looking Ahead
By shifting from direct HTTP payload transfers to a decoupled GCS storage and streaming model, MedReachAI will be fully equipped to handle enterprise-grade healthcare datasets smoothly and reliably.

Stay tuned as Collin and I execute these Sprint 6 tasks and bring our capstone architecture to its production-ready finish line!

Top comments (0)