Ovis
by AIDC-AI
Pythonpushed almost 2 years ago
A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and textual embeddings.
AI summary
Multimodal aligner
An MLLM architecture designed to align visual and textual embeddings through structural alignment
- stars
- 575
- forks
- 33
- watching
- 7