Sight-Beyond-Text

Multimodal model trainer

An implementation of a multimodal LLM training paradigm to enhance truthfulness and ethics in language models

[TMLR 2024] Official implementation of "Sight Beyond Text: Multi-Modal Training Enhances LLMs in Truthfulness and Ethics"

GitHub

19 stars
2 watching
1 forks
Language: Python
last commit: about 3 years ago
ai-alignmentalignmentllama2llavallmmllmvicunavision-languagevlm

Related projects:

RepositoryDescriptionStars
mlpc-ucsd/blivaA multimodal LLM designed to handle text-rich visual questions270
ucsc-vlaa/vllm-safety-benchmarkA benchmark for evaluating the safety and robustness of vision language models against adversarial attacks.72
vpgtrans/vpgtransTransfers visual prompt generators across large language models to reduce training costs and enable customization of multimodal LLMs270
ailab-cvc/seedAn implementation of a multimodal language model with capabilities for comprehension and generation585
lyuchenyang/macaw-llmA multi-modal language model that integrates image, video, audio, and text data to improve language understanding and generation1,568
aidc-ai/ovisAn MLLM architecture designed to align visual and textual embeddings through structural alignment575
mbzuai-oryx/groundinglmmAn end-to-end trained model capable of generating natural language responses integrated with object segmentation masks for interactive visual conversations797
pleisto/yuren-baichuan-7bA multi-modal large language model that integrates natural language and visual capabilities with fine-tuning for various tasks73
llava-vl/llava-plus-codebaseA platform for training and deploying large language and vision models that can use tools to perform tasks717
vishaal27/sus-xThis is an open-source project that proposes a novel method to train large-scale vision-language models with minimal resources and no fine-tuning required.94
csuhan/onellmA framework for training and fine-tuning multimodal language models on various data types601
salt-nlp/llavarAn open-source project that enhances visual instruction tuning for text-rich image understanding by integrating GPT-4 models with multimodal datasets.259
alpha-vllm/wemix-llmAn LLaMA-based multimodal language model with various instruction-following and multimodal variants.17
bobazooba/xllmA tool for training and fine-tuning large language models using advanced techniques387
neulab/pangeaAn open-source multilingual large language model designed to understand and generate content across diverse languages and cultural contexts92