StackOverflow-Question-Code-Dataset

Question dataset

A collection of mined question-code pairs from Stack Overflow used for training and testing AI models

StaQC: a systematically mined dataset containing around 148K Python and 120K SQL domain question-code pairs, as described in "StaQC: A Systematically Mined Question-Code Dataset from Stack Overflow" (WWW'18)

GitHub

166 stars
7 watching
28 forks
Language: Python
last commit: about 5 years ago
Linked from 1 awesome list


Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
crownpku/small-chinese-corpusA collection of datasets and tools for NLP tasks on Chinese texts, including part-of-speech tagging, named entity recognition, and question answering.529
ysu1989/graphquestionsA characteristic-rich dataset for factoid question answering with explicit specification of question characteristics and logical forms.92
src-d/datasetsProvides datasets and tools for analyzing source code in various aspects such as programming languages, commits, and more.323
ymcui/cmrc2018A collection of data for evaluating Chinese machine reading comprehension systems419
pku-yuangroup/open-sora-datasetA large video dataset collected from various open-source websites for use in computer vision and multimedia applications.94
certainlyio/corona_datasetA collection of data to train chatbots on COVID-19-related questions11
maluuba/newsqaCompiles and provides structured access to Maluuba's NewsQA dataset for natural language question answering research.253
thu-coai/cdial-gptA large-scale Chinese conversation dataset and pre-trained dialog models for text generation1,799
srush/minichainA tiny library for using large language models in code generation and debugging1,221
ujjwalkarn/datasciencepythonA curated list of tutorials and resources for learning Python for data science, machine learning, and other related topics.5,301
websail-nu/codahReleases an adversarially constructed commonsense question-answering dataset for testing common sense in natural language understanding22
fido-ai/ua-datasetsProvides a collection of datasets for natural language processing in Ukrainian.57
witiko/semeval-2016_2017-task3-subtaskb-englishConverts XML datasets to JSON for community question answering task 3 subtask b1
pratyushmaini/llm_dataset_inferenceDetects whether a given text sequence is part of the training data used to train a large language model.23
mikegu721/xiezhibenchmarkAn evaluation suite to assess language models' performance in multi-choice questions93