TensorRT-LLM
by NVIDIA
C++pushed almost 2 years ago
TensorRT-LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and build TensorRT engines that contain state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT-LLM also contains components to create Python and C++ runtimes that execute those TensorRT engines.
AI summary
Inference optimizer
A software framework providing an easy-to-use Python API to optimize Large Language Models on NVIDIA GPUs for efficient inference.
- stars
- 8.9K
- forks
- 1K
- watching
- 95
- awesome list
- 1
Featured in 1 awesome list
Each link jumps to the spot where the list mentions TensorRT-LLM.