I open-sourced Learn TensorRT, a hands-on course for developers moving from PyTorch models to C++ deployment. The baseline is TensorRT 10.14, CUDA 13.0 and C++17, in a pinned NVIDIA development container.
The core path uses YOLOv8 to connect three parts of deployment:
- Correctness: export to ONNX, inspect outputs with Polygraphy, then implement preprocessing, TensorRT inference and postprocessing in C++.
- Optimization: compare FP32/FP16/INT8, use explicit Q/DQ quantization, and profile with Nsight.
- Pipeline behavior: add bounded queues, dynamic batching and asynchronous CUDA streams, then measure latency, throughput and overload behavior.
A useful starting point is the single-image C++ pipeline in lesson 11. Follow its prerequisites, check the detections, and establish correctness before changing precision or adding concurrency. Later reports separate engine timing from application and pipeline measurements.
The code emphasizes RAII, explicit resource ownership and target-based CMake. Lessons have build/run instructions and reporting checkpoints. Ubuntu with an NVIDIA GPU is the reference setup; C++ and CMake basics are expected. Engines and performance results must be regenerated for your environment.
Electives include plugins, Triton, DeepStream and Jetson/DLA. Some elective runtime acceptance is still pending; the coverage matrix records those limits.
Repository and learning roadmap — MIT licensed, with English, Chinese, Japanese and Korean READMEs.
If you're learning TensorRT deployment, try the core path. Feedback on unclear steps or reproducibility problems is welcome.
AI assistance was used to draft this introduction.
Top comments (1)
Putting single-image correctness before precision and concurrency gives lesson 11 a useful baseline. The async extension introduces an ownership boundary that RAII alone does not settle: a host scope can end while a CUDA stream still uses its buffers, or a pool can reuse an output slot before its completion event fires.
A focused exercise would submit two distinct images on nondefault streams, deliberately stagger completion, and verify that each result remains attached to its original request. Buffer leases should survive until the relevant event completes. That would connect the resource-ownership lessons to a concurrency bug that single-image output comparisons cannot expose.