VL-ICL
by ys-zong
Code for paper: VL-ICL Bench: The Devil in the Details of Benchmarking Multimodal In-Context Learning
AI summary
Learning benchmark
A benchmarking suite for multimodal in-context learning models
- stars
- 31
- forks
- 2
- watching
- 1
Similar projects
Found by comparing what the projects do, not just their names.
Multimodal model evaluator
Evaluating and improving large multimodal models through in-context learning
ML Model Benchmarker
Provides a benchmarking framework and dataset for evaluating the performance of large language models in text-to-image tasks
Safety fine-tuner
Improves safety and helpfulness of large language models by fine-tuning them using safety-critical tasks
Multimodal LLM test suite
A benchmark for evaluating large language models' ability to process multimodal input
RL benchmarks
A collection of benchmarks and implementations for testing reinforcement learning-based Volt-VAR control algorithms
Multimodal learner
Develops a multimodal vision-language model to enable machines to understand complex relationships between instructions and images in various tasks.
OCR Benchmark
An evaluation benchmark for OCR capabilities in large multmodal models.
Multimodal benchmarking
Evaluates and benchmarks multimodal language models' ability to process visual, acoustic, and textual inputs simultaneously.
ydli-ai/csl582
Chinese Scientific Dataset
A large-scale dataset for natural language processing tasks focused on Chinese scientific literature, providing tools and benchmarks for NLP research.
cloud-cv/evalai1.8K
Benchmarking tool
A platform for comparing and evaluating AI and machine learning algorithms at scale
RL Benchmark
A benchmark suite for unsupervised reinforcement learning agents, providing pre-trained models and scripts for testing and fine-tuning agent performance.
ML models
Provides pre-trained machine learning models for natural language processing tasks using Clojure and the clj-djl framework.
Vision-Language Learning Model
Develops and trains models for vision-language learning with decoupled language pre-training
Visual Knowledge Model
This project integrates visual knowledge into large language models to improve their capabilities and reduce hallucinations.
Multimodal alignment model
Extending pretraining models to handle multiple modalities by aligning language and video representations