VinVL

Visual representation improvement

A project aimed at improving visual representations in vision-language models by developing an object detection model for richer visual object and concept representations.

project page for VinVL

GitHub

350 stars
9 watching
25 forks
last commit: about 3 years ago

Related projects:

RepositoryDescriptionStars
zuoxingdong/vin_pytorch_visdomImplementation of Value Iteration Networks in PyTorch with visualization capabilities using Visdom.226
dotnet/vblangDesign of a Visual Basic .NET language and runtime library291
yunxinli/lingcloudEnhances language models by incorporating human-like eyes to improve visual comprehension and interaction with external world48
lucasvazq/lucasvazqPersonal showcase of a developer's experience and expertise in building web platforms with a focus on accessibility, performance, and robust code.30
pku-yuangroup/chat-univiA framework for unified visual representation in image and video understanding models, enabling efficient training of large language models on multimodal data.895
ivanreese/visual-programming-codexAn online resource showcasing alternative visual programming languages and projects, reflecting on their concepts, ideas, and potential benefits.1,366
byungkwanlee/collavoDevelops a PyTorch implementation of an enhanced vision language model93
liaoning97/revo-lionA comprehensive dataset and evaluation framework for Vision-Language Instruction Tuning models11
vanshkapoor/vanshkapoorUtility tools and project showcases24
dbuenzli/vgA declarative 2D vector graphics library written in OCaml91
byungkwanlee/moaiImproves performance of vision language tasks by integrating computer vision capabilities into large language models314
nickjiang2378/vl-interpThis project provides an official PyTorch implementation of a method to interpret and edit vision-language representations to mitigate hallucinations in image captions.46
ys-zong/vl-iclA benchmarking suite for multimodal in-context learning models31
kunpengli1994/vsrnAn open-source PyTorch implementation of a visual semantic reasoning model for image-text matching294
brucesherwood/vpython-jupyterAn integration of VPython with Jupyter Notebook for interactive 3D visualization and simulation in scientific computing.64