scenic

Computer Vision Toolkit

A collection of libraries and projects focused on research around attention-based models for computer vision and beyond, providing optimized tools and baselines.

Scenic: A Jax Library for Computer Vision Research and Beyond

GitHub

3k stars
39 watching
444 forks
Language: Python
last commit: almost 2 years ago
Linked from 1 awesome list

attentioncomputer-visiondeep-learningjaxresearchtransformersvision-transformer

Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
google-research/vision_transformerProvides pre-trained models and code for training vision transformers and mixers using JAX/Flax10,620
roboflow/notebooksThis repository contains tutorials and examples on using state-of-the-art computer vision models and techniques5,678
google-research/big_visionSupports large-scale vision model training on GPU machines or Google Cloud TPUs using scalable input pipelines.2,439
jshilong/gpt4roiTraining and deploying large language models on computer vision tasks using region-of-interest inputs517
uber-research/upsnetDevelops an instance segmentation and panoptic segmentation model for computer vision tasks.648
haotian-liu/llavaA system that uses large language and vision models to generate and process visual instructions20,683
google-research/cad-estateA large dataset of 3D object and room layout annotations on RGB videos, designed to test automatic scene understanding methods.106
google-research/big_transferPre-trained models and code for fine-tuning image recognition tasks using deep learning frameworks1,516
google/jaxoptAn open-source project providing hardware accelerated, batchable and differentiable optimizers in JAX for deep learning.941
nexusapoorvacus/deepvariationstructuredrlAn implementation of reinforcement learning for visual relationship and attribute detection using PyTorch.63
huggingface/transformersA collection of pre-trained machine learning models for various natural language and computer vision tasks, enabling developers to fine-tune and deploy these models on their own projects.136,357
deepseek-ai/deepseek-vlA multimodal AI model that enables real-world vision-language understanding applications2,145
rastapasta/react-native-gl-model-viewA React Native component that displays and animates 3D models loaded from Wavefront OBJ files.419
vision-cair/minigpt-4Enabling vision-language understanding by fine-tuning large language models on visual data.25,490
matthias-wright/flaxmodelsProvides pre-trained deep learning models for the Jax/Flax ecosystem.240