ThermaCompute AI is building diagnostics to help AI infrastructure teams understand GPU failures and performance waste. Our current focus is passive investigation: useful evidence and recommendations an operator can test.
A final timeout can hide the earlier error that deserves attention. Before restarting a failed job, I would work through this checklist.
1. Reconstruct the sequence
Record the job, worker or rank, timestamps and source log. Check whether a worker error occurred before the collective timeout. An ordering is a lead to investigate, not proof that one event caused the other.
2. Separate temperature from throttling
A temperature reading alone does not prove thermal throttling. Look for supported throttle-reason telemetry, clock changes and workload throughput in the same time window. Power limiting, memory pressure and communication failures need separate consideration. Missing telemetry should remain an explicit unknown.
3. Make the next check reproducible
A useful report should include the raw evidence line, the rule that matched, what the finding does not establish, and a concrete next check. Avoid turning a log match into an exact dollar-loss estimate without a measured baseline and stated cost assumptions.
Try the free CrashLens example
We made CrashLens to turn a saved log into a focused investigation. It is an early, open-source CLI for supported OOM, NCCL and NVIDIA Xid patterns, with evidence lines and suggested checks. It does not control GPUs or establish root cause.
The synthetic demo needs Python 3.10+, no GPU or account, and no log upload:
git clone https://github.com/rohitcn-hub/thermacompute-crashlens.git
cd thermacompute-crashlens
python crashlens.py examples/synthetic.log --out demo-report
Open demo-report/report.html in your browser. Choose a new output directory if it already exists. Preview the synthetic report before installing.
For GPU operators: did the report give you one useful next check? What evidence would be missing before you would test it? Please do not post private production logs.
Optional scoped review
ThermaCompute offers an $800 Executive Thermal Architecture Audit for teams with suitable NVML/vLLM telemetry. We agree one workload/cohort and time window, then deliver a PDF with supported findings, limitations and prioritized recommendations for testing. Data handling and scope are agreed before transfer; savings and exhaustive fault detection are not guaranteed. Email vivaan.thermacompute@gmail.com with “Executive audit” to discuss fit. CrashLens remains free regardless.
Disclosure: I am the founder of ThermaCompute and maintain CrashLens. This article was prepared with AI assistance.
Top comments (0)