Pink

Visual Text Understanding

This project enables multi-modal language models to understand and generate text about visual content using referential comprehension.

Pink: Unveiling the Power of Referential Comprehension for Multi-modal LLMs

GitHub

79 stars
4 watching
5 forks
Language: Python
last commit: over 2 years ago

Related projects:

RepositoryDescriptionStars
airaria/visual-chinese-llama-alpacaDevelops a multimodal Chinese language model with visual capabilities429
lxtgh/omg-segDevelops an end-to-end model for multiple visual perception and reasoning tasks using a single encoder, decoder, and large language model.1,336
ys-zong/vl-iclA benchmarking suite for multimodal in-context learning models31
sergioburdisso/pyss3A Python package implementing an interpretable machine learning model for text classification with visualization tools336
yuliang-liu/monkeyAn end-to-end image captioning system that uses large multi-modal models and provides tools for training, inference, and demo usage.1,849
pku-yuangroup/languagebindExtending pretraining models to handle multiple modalities by aligning language and video representations751
jiutian-vl/jiutian-lionThis project integrates visual knowledge into large language models to improve their capabilities and reduce hallucinations.124
dvlab-research/prompt-highlighterAn interactive control system for text generation in multi-modal language models135
pleisto/yuren-baichuan-7bA multi-modal large language model that integrates natural language and visual capabilities with fine-tuning for various tasks73
yfzhang114/slimeDevelops large multimodal models for high-resolution understanding and analysis of text, images, and other data types.143
brightmart/xlnet_zhTrains a large Chinese language model on massive data and provides a pre-trained model for downstream tasks230
yiren-jian/blitextDevelops and trains models for vision-language learning with decoupled language pre-training24
yunxinli/lingcloudEnhances language models by incorporating human-like eyes to improve visual comprehension and interaction with external world48
penghao-wu/vstarPyTorch implementation of guided visual search mechanism for multimodal LLMs541
m-clark/visiblyA collection of R visualization tools and utilities for creating color palettes, themes, and visualizing statistical models.63