deepsparse

Inference runtime

A sparsity-aware deep learning inference runtime for CPUs that optimizes neural network performance on CPU hardware.

Sparsity-aware deep learning inference runtime for CPUs

GitHub

3k stars
57 watching
175 forks
Language: Python
last commit: about 2 years ago
Linked from 1 awesome list

computer-visioncpusdeepsparseinferencellm-inferencemachinelearningnlpobject-detectiononnxperformancepretrained-modelspruningquantizationsparsification

Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
neuralmagic/sparsemlEnables the creation of smaller neural network models through efficient pruning and quantization techniques2,083
intel/neural-compressorTools and techniques for optimizing large language models on various frameworks and hardware platforms.2,257
microsoft/deepspeedA deep learning optimization library that simplifies distributed training and inference on modern computing hardware.35,863
mit-han-lab/llm-awqAn open-source software project that enables efficient and accurate low-bit weight quantization for large language models.2,593
labmlai/annotated_deep_learning_paper_implementationsImplementations of various deep learning algorithms and techniques with accompanying documentation57,177
confident-ai/deepevalA framework for evaluating large language models4,003
ludwig-ai/ludwigA low-code framework for building custom deep learning models and neural networks11,236
oxford-cs-deepnlp-2017/lecturesAn open-source repository containing lecture slides and course materials for an advanced natural language processing course.15,702
deepseek-ai/deepseek-v2A high-performance mixture-of-experts language model with strong performance and efficient inference capabilities.3,758
fminference/flexllmgenGenerates large language model outputs in high-throughput mode on single GPUs9,236
google-deepmind/deepmind-researchProvides implementations and illustrative code to accompany DeepMind research publications13,329
microsoft/lightgbmA high-performance gradient boosting framework for machine learning tasks16,769
modeltc/lightllmA Python-based framework for serving large language models with low latency and high scalability.2,691
optimalscale/lmflowA toolkit for fine-tuning and inferring large machine learning models8,312
lyogavin/airllmOptimizes large language model inference on limited GPU resources5,446