DeepSpeed-MII

Inference accelerator

A Python library designed to accelerate model inference with high-throughput and low latency capabilities

MII makes low-latency and high-throughput inference possible, powered by DeepSpeed.

GitHub

2k stars
42 watching
175 forks
Language: Python
last commit: almost 2 years ago
Linked from 2 awesome lists

deep-learninginferencepytorch

Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
utensor/utensorA lightweight machine learning inference framework built on Tensorflow optimized for Arm targets.1,742
xboot/libonnxAn onnx inference engine for embedded devices with hardware acceleration support589
megvii-research/tlcImproves image restoration performance by converting global operations to local ones during inference231
xilinx/finnFast and scalable neural network inference framework for FPGAs.770
microsoft/archaiAutomates the search for optimal neural network configurations in deep learning applications468
microsoft/megatron-deepspeedResearch tool for training large transformer language models at scale1,926
mims-harvard/ohmnetAn algorithm for learning feature representations in multi-layer networks81
torchpipe/torchpipeAn open-source framework that enables the deployment and serving of PyTorch models in various acceleration frameworks.147
jgreenemi/parrisAutomates the setup and training of machine learning algorithms on remote servers316
lge-arc-advancedai/auptimizerAutomates model building and deployment process by optimizing hyperparameters and compressing models for edge computing.200
denosaurs/netsaurA machine learning library with GPU, CPU, and WASM backends for building neural networks.233
arthurpaulino/miraimlAn asynchronous engine for continuous and autonomous machine learning26
intel/neural-compressorTools and techniques for optimizing large language models on various frameworks and hardware platforms.2,257
mlcommons/inferenceMeasures the performance of deep learning models in various deployment scenarios.1,256