XVERSE-V-13B
by xverse-ai
AI summary
Multimodal model
A large multimodal model for visual question answering, trained on a dataset of 2.1B image-text pairs and 8.2M instruction sequences.
- stars
- 78
- forks
- 4
- watching
- 4
Similar projects
Found by comparing what the projects do, not just their names.
Language Model
A large language model developed to support multiple languages and applications
Multilingual Model
Develops and publishes large multilingual language models with advanced mixing-of-experts architecture.
Mixture-of-Experts Model
Developed by XVERSE Technology Inc. as a multilingual large language model with a unique mixture-of-experts architecture and fine-tuned for various tasks such as conversation, question answering, and natural language understanding.
Language Model
A large language model developed by XVERSE Technology Inc. using transformer architecture and fine-tuned on diverse data sets for various applications.
LLM
A multilingual large language model developed by XVERSE Technology Inc.
Multimodal model developer
Develops large multimodal models for high-resolution understanding and analysis of text, images, and other data types.
Multimodal evaluation framework
Develops a multimodal task and dataset to assess vision-language models' ability to handle interleaved image-text inputs.
openbmb/viscpm1.1K
Multimodal Models
A family of large multimodal models supporting multimodal conversational capabilities and text-to-image generation in multiple languages
tsb0601/mmvp296
Visual model evaluation
An evaluation framework for multimodal language models' visual capabilities using image and question benchmarks.
Retrieval model trainer
Trains and evaluates a universal multimodal retrieval model to perform various information retrieval tasks.
Model arena
An evaluation platform for comparing multi-modality models on visual question-answering tasks
nvlabs/eagle549
Multimodal model builder
Develops high-resolution multimodal LLMs by combining vision encoders and various input resolutions
VQA model
A multimodal LLM designed to handle text-rich visual questions
Robot learner
An implementation of a general-purpose robot learning model using multimodal prompts
Multimodal benchmarking
Evaluates and benchmarks multimodal language models' ability to process visual, acoustic, and textual inputs simultaneously.