Huatuo-26M
The Largest-scale Chinese Medical QA Dataset: with 26,000,000 question answer pairs.
AI summary
Medical QA dataset
A large-scale medical question-and-answer dataset with over 26 million high-quality pairs, designed for natural language processing and machine learning applications in the medical field.
- stars
- 226
- forks
- 22
- watching
- 9
Similar projects
Found by comparing what the projects do, not just their names.
Medical LLM
Developing a large language model for medical consultations by combining distilled and real-world data to improve doctor-patient interactions
Medical info model
A large-scale Chinese medical language model trained on diverse data sources to provide accurate and reliable medical information
Medical chatbot
Develops a large language model capable of handling complex medical conversations with high accuracy and professionalism.
Covid Q&A dataset
A collection of data to train chatbots on COVID-19-related questions
Medical image understanding toolkit
A medical visual question-answering dataset and toolkit for training models to understand medical images and instructions.
News QA data
Compiles and provides structured access to Maluuba's NewsQA dataset for natural language question answering research.
Instruction Dataset
Creating a large-scale user-based instruction dataset for natural language processing research and development
Vision-Language Model Dataset
A collection of datasets and models designed to support the training of lite vision-language models.
Medical diagnosis assistant
Develops a large language model to aid in Chinese medicine diagnosis and prescription recommendations.
Image QA model
This project provides code for training image question answering models using stacked attention networks and convolutional neural networks.
Medical NLP training data
A large-scale instruction-tuning dataset for multi-task and few-shot learning in the medical domain
Face dataset
Provides images of human faces with annotated age, gender, pose, and other attributes for testing age transformation algorithms.
Insurance dataset
An insurance industry conversation corpus with pre-processed data for natural language processing and question answering tasks.
Medical dialogue dataset
A collection of medical dialogue data for training conversational AI models.
Quantum datasets
A collection of 52 machine learning datasets for simulating quantum systems with noise and controls.