Fair-LLM-Benchmark

Bias datasets

Compiles bias evaluation datasets and provides access to original data sources for large language models

GitHub

115 stars
4 watching
9 forks
Language: Python
last commit: about 3 years ago

Related projects:

RepositoryDescriptionStars
dssg/aequitasToolkit to audit and mitigate biases in machine learning models701
freedomintelligence/mllm-benchEvaluates and compares the performance of multimodal large language models on various tasks56
damo-nlp-sg/m3examA benchmark for evaluating large language models in multiple languages and formats93
privacytrustlab/bias_in_flThis project investigates how bias can be introduced and spread in machine learning models during federated learning, and aims to detect and mitigate this issue.11
mlgroupjlu/llm-eval-surveyA repository of papers and resources for evaluating large language models.1,450
aifeg/benchlmmAn open-source benchmarking framework for evaluating cross-style visual capability of large multimodal models84
nyu-mll/bbqA dataset and benchmarking framework to evaluate the performance of question answering models on detecting and mitigating social biases.92
mbilalzafar/fair-classificationProvides a Python implementation of fairness mechanisms in classification models to mitigate disparate impact and mistreatment.190
qcri/llmebenchA benchmarking framework for large language models81
modeloriented/fairmodelsA tool for detecting bias in machine learning models and mitigating it using various techniques.86
ethicalml/xaiAn eXplainability toolbox for machine learning that enables data analysis and model evaluation to mitigate biases and improve performance1,135
btschwertfeger/python-cmethodsA collection of bias correction techniques for climate data analysis60
adebayoj/fairmlAn auditing toolbox to assess the fairness of black-box predictive models361
google/ml-fairness-gymAn open-source tool for simulating the long-term impacts of machine learning-based decision systems on social environments314
cisco-open/inclusive-languageTools and resources for identifying biased language in code and content.21