VIMA

Robot learner

An implementation of a general-purpose robot learning model using multimodal prompts

Official Algorithm Implementation of ICML'23 Paper "VIMA: General Robot Manipulation with Multimodal Prompts"

GitHub

781 stars
17 watching
88 forks
Language: Python
last commit: over 2 years ago

Related projects:

RepositoryDescriptionStars
ailab-cvc/seedAn implementation of a multimodal language model with capabilities for comprehension and generation585
autoviml/auto_vimlAutomatically builds multiple machine learning models using a single line of code.526
jdelacroix/simiamEducational tool for robotics that bridges theory and practice using MATLAB103
vita-epfl/crowdnavDevelops robot navigation policies in crowded spaces using reinforcement learning and attention mechanisms.607
open-mmlab/multimodal-gptTrains a multimodal chatbot that combines visual and language instructions to generate responses1,478
llava-vl/llava-interactive-demoAn all-in-one demo for interactive image processing and generation353
xverse-ai/xverse-v-13bA large multimodal model for visual question answering, trained on a dataset of 2.1B image-text pairs and 8.2M instruction sequences.78
sergioburdisso/pyss3A Python package implementing an interpretable machine learning model for text classification with visualization tools336
dvlab-research/prompt-highlighterAn interactive control system for text generation in multi-modal language models135
vpgtrans/vpgtransTransfers visual prompt generators across large language models to reduce training costs and enable customization of multimodal LLMs270
multimodal-art-projection/omnibenchEvaluates and benchmarks multimodal language models' ability to process visual, acoustic, and textual inputs simultaneously.15
ethereum/evmlabUtilities for interacting with the Ethereum virtual machine367
bekovmi/segmentation_tutorialA tutorial project on teaching model training scripts using Config API Catalyst8
zhourax/vegaDevelops a multimodal task and dataset to assess vision-language models' ability to handle interleaved image-text inputs.33
mkorpela/robomachineAutomates test generation based on user input and system behavior models.101