Awesome Lists

TensorRT-LLM

by NVIDIA

C++pushed almost 2 years ago

TensorRT-LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and build TensorRT engines that contain state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT-LLM also contains components to create Python and C++ runtimes that execute those TensorRT engines.

AI summary

Inference optimizer

A software framework providing an easy-to-use Python API to optimize Large Language Models on NVIDIA GPUs for efficient inference.

stars
8.9K
forks
1K
watching
95
awesome list
1
View on GitHubnvidia.github.io/TensorRT-LLM

Embed the badge

Show how many awesome lists link to your project. The count updates automatically.

Awesome Lists badge
Markdown
[![Awesome Lists Badge](https://awesome.facts.dev/shield/NVIDIA/TensorRT-LLM/links.svg)](https://awesome.facts.dev/awesome/NVIDIA/TensorRT-LLM)
HTML
<a href="https://awesome.facts.dev/awesome/NVIDIA/TensorRT-LLM"><img src="https://awesome.facts.dev/shield/NVIDIA/TensorRT-LLM/links.svg" alt="Awesome Lists Badge" /></a>
Image URL
https://awesome.facts.dev/shield/NVIDIA/TensorRT-LLM/links.svg

Add a GitHub project

Missing a project or an awesome list? Paste its GitHub URL and we fetch it right away.