GPT4Tools

Conversational image interface

An intelligent system that enables automatic control and utilization of visual foundation models to interact with images in conversational settings.

GPT4Tools is an intelligent system that can automatically decide, control, and utilize different visual foundation models, allowing the user to interact with images during a conversation.

GitHub

762 stars
13 watching
59 forks
Language: Python
last commit: over 2 years ago
Linked from 1 awesome list


Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
jshilong/gpt4roiTraining and deploying large language models on computer vision tasks using region-of-interest inputs517
vinhnx/inkchatgptAn application that enables users to upload documents and converse with an AI-powered language model.9
thu-coai/cdial-gptA large-scale Chinese conversation dataset and pre-trained dialog models for text generation1,799
open-mmlab/multimodal-gptTrains a multimodal chatbot that combines visual and language instructions to generate responses1,478
neukg/techgpt-2.0An advanced language model designed to generate human-like responses in various domains and applications101
fengyuli-dev/multimedia-gptEnables OpenAI GPT to process multimedia inputs like images and audio with text output184
chidiwilliams/gpt-automatorA voice-controlled Mac assistant that uses natural language processing to automate desktop tasks232
robitx/gp.nvimAn extension for Neovim that integrates GPT models into the editor, enabling AI-powered text operations and speech-to-text capabilities.928
mbzuai-oryx/video-chatgptA video conversation model that generates meaningful conversations about videos using large vision and language models1,246
kejunmao/ai-anythingAn open-source toolset for creating custom ChatGPT interfaces568
zcli-charlie/batgptA large language model designed to support long context conversations with improved efficiency and effectiveness38
360cvgroup/seechatA multimodal chatbot with computer vision capabilities integrated into a single model99
pjlab-adg/gpt4v-ad-explorationAn autonomous driving project exploring the capabilities of a visual-language model in understanding complex driving scenes and making decisions288
laurentkneip/opengvA collection of computer vision methods for solving geometric vision problems1,040
ailab-cvc/seed-benchA benchmark for evaluating large language models' ability to process multimodal input322