LingCloud
by YunxinLi
Attaching human-like eyes to the large language model. The codes of IEEE TMM paper "LMEye: An Interactive Perception Network for Large Language Model""
AI summary
Visual enhancer for LLMs
Enhances language models by incorporating human-like eyes to improve visual comprehension and interaction with external world
- stars
- 48
- forks
- 1
- watching
- 1
Similar projects
Found by comparing what the projects do, not just their names.
Visual Knowledge Model
This project integrates visual knowledge into large language models to improve their capabilities and reduce hallucinations.
MIL booster
An open-source software package implementing two boosting-based Multiple Instance Learning methods for image segmentation and classification tasks.
Visual representation improvement
A project aimed at improving visual representations in vision-language models by developing an object detection model for richer visual object and concept representations.
Vision-Language Learning Model
Develops and trains models for vision-language learning with decoupled language pre-training
Visual Text Understanding
This project enables multi-modal language models to understand and generate text about visual content using referential comprehension.
Token representation refinement
Improves pre-trained language models by encouraging an isotropic and discriminative distribution of token representations.
Vision language model trainer
An annotated preference dataset and training framework for improving large vision language models.
Learning booster
A semi-supervised learning method to improve the accuracy of machine learning models by using noisy teacher models and student models.
Perception adapter
An adapter for improving large language models at object-level perception tasks with auxiliary perception modalities
Visualization library
A collection of React components for creating animated and interactive visualizations.
Boosting algorithm
An implementation of online multi-label ranking boosting using VFDT as weak learners
Vision-Language Model Trainer
This is an open-source project that proposes a novel method to train large-scale vision-language models with minimal resources and no fine-tuning required.
Benchmark
An image-context reasoning benchmark designed to challenge large vision-language models and help improve their accuracy
Chinese NLU/NGL toolkit
This project provides pre-trained models and tools for natural language understanding (NLU) and generation (NLG) tasks in Chinese.
Model validator
Analyzing and mitigating object hallucination in large vision-language models to improve their accuracy and reliability.