flash-attention

Attention algorithms

Implementations of efficient exact attention mechanisms for machine learning

Fast and memory-efficient exact attention

GitHub

15k stars
122 watching
1k forks
Language: Python
last commit: almost 2 years ago
Linked from 1 awesome list


Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
facebookincubator/aitemplateA framework that transforms deep neural networks into high-performance GPU-optimized C++ code for efficient inference serving.4,573
luolc/adaboundAn optimizer that combines the benefits of Adam and SGD algorithms2,908
facebookresearch/slowfastProvides state-of-the-art video understanding codebase with efficient training methods and pre-trained models for various tasks6,680
albumentations-team/albumentationsA Python library providing a flexible and fast image augmentation tool for machine learning and computer vision tasks.14,386
arrayfire/arrayfireA high-level abstraction of data on parallel architectures for efficient tensor computing and machine learning applications.4,587
microsoft/flamlAutomates machine learning workflows and optimizes model performance using large language models and efficient algorithms3,968
rapidsai/cudfA GPU-accelerated data manipulation library built on top of C++/CUDA and Apache Arrow.8,534
dynamorio/drmemoryAn open-source memory debugger for multiple operating systems and platforms2,468
bytedance/bytepsA high-performance distributed deep learning framework supporting multiple frameworks and networks3,635
rapidsai/cumlA suite of libraries implementing machine learning algorithms and mathematical primitives on NVIDIA GPUs4,292
huggingface/accelerateA tool to simplify training and deployment of PyTorch models on various devices and configurations8,056
microsoft/deepspeedA deep learning optimization library that simplifies distributed training and inference on modern computing hardware.35,863
flashlight/flashlightA C++ machine learning library with autograd support and high-performance defaults for efficient computation.5,300
ntop/pf_ringA framework for high-speed packet processing on Linux kernels.2,718
tencent/pocketflowA framework that automatically compresses and accelerates deep learning models to make them suitable for mobile devices with limited computational resources.2,787