flash-attention
by Dao-AILab
Fast and memory-efficient exact attention
AI summary
Attention algorithms
Implementations of efficient exact attention mechanisms for machine learning
- stars
- 14.7K
- forks
- 1.4K
- watching
- 122
- awesome list
- 1
Similar projects
Found by comparing what the projects do, not just their names.
Neural network optimizer
A framework that transforms deep neural networks into high-performance GPU-optimized C++ code for efficient inference serving.
luolc/adabound2.9K
Optimizer
An optimizer that combines the benefits of Adam and SGD algorithms
Video understanding library
Provides state-of-the-art video understanding codebase with efficient training methods and pre-trained models for various tasks
Image Augmentation Library
A Python library providing a flexible and fast image augmentation tool for machine learning and computer vision tasks.
Tensor library
A high-level abstraction of data on parallel architectures for efficient tensor computing and machine learning applications.
Machine learning automator
Automates machine learning workflows and optimizes model performance using large language models and efficient algorithms
rapidsai/cudf8.5K
GPU DataFrame Library
A GPU-accelerated data manipulation library built on top of C++/CUDA and Apache Arrow.
Memory debugger
An open-source memory debugger for multiple operating systems and platforms
bytedance/byteps3.6K
Distributed DL framework
A high-performance distributed deep learning framework supporting multiple frameworks and networks
rapidsai/cuml4.3K
GPU ML library
A suite of libraries implementing machine learning algorithms and mathematical primitives on NVIDIA GPUs
Model trainer
A tool to simplify training and deployment of PyTorch models on various devices and configurations
microsoft/deepspeed35.9K
Deep Learning Optimizer
A deep learning optimization library that simplifies distributed training and inference on modern computing hardware.
ML library
A C++ machine learning library with autograd support and high-performance defaults for efficient computation.
ntop/pf_ring2.7K
Packet processor
A framework for high-speed packet processing on Linux kernels.
Model compressor
A framework that automatically compresses and accelerates deep learning models to make them suitable for mobile devices with limited computational resources.
Featured in 1 awesome list
Each link jumps to the spot where the list mentions flash-attention.