Hi everyone,
We're three final-year students at Epitech building Horus, a multi-agent assistant that runs entirely locally and offline. Our current challenge is hardware: keeping it usable on machines without a powerful GPU, without long setup times, excessive RAM/VRAM use or crashes.
Where we are today, from our last beta test:
One tester needed over 2 hours to install. The Python dependencies alone take 21–46 min, and the download is about 25 GB.
On CPU only, routing a question can take around 40 s, long enough for our WebSocket connection to drop.
[Models we use + the smallest machine we've tested on]
We'd love advice from anyone experienced with:
CPU-only LLM inference and memory-efficient loading
Quantization and model choice for low-end hardware
GPU/CPU fallback strategies
Hardware detection and adaptive configuration
Preventing resource exhaustion during setup and execution
Advice in the comments is very welcome, with no strings attached.
Looking for contributors: we also have a few small, well-scoped tasks or code reviews (about 1–4 hours), for example [reviewing our hardware detection and model selection, or benchmarking a quantized model on a 16 GB RAM laptop].
To be transparent: Horus is closed source and will be licensed to companies. Contributing is voluntary and unpaid. Before seeing any code, contributors sign a short confidentiality and contributor agreement, and the code they contribute becomes part of Horus. In return we offer thorough code reviews, full credit in the project and a professional reference on request.
We're not sharing code or private links publicly. If this interests you, comment below or DM me with your experience in local inference, CPU optimisation or offline apps
Top comments (0)