TrustLLM

Trust Assessment Toolkit

A toolkit for assessing trustworthiness in large language models

[ICML 2024] TrustLLM: Trustworthiness in Large Language Models

GitHub

491 stars
8 watching
47 forks
Language: Python
last commit: almost 2 years ago
Linked from 1 awesome list

aibenchmarkdatasetevaluationlarge-language-modelsllmnatural-language-processingnlppypi-packagetoolkittrustworthy-aitrustworthy-machine-learning

Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
aiplanethub/beyondllmAn open-source toolkit for building and evaluating large language models267
vhellendoorn/code-lmsA guide to using pre-trained large language models in source code analysis and generation1,789
leondz/lm_risk_cardsA set of tools and guidelines for assessing the security vulnerabilities of language models in AI applications28
mlgroupjlu/llm-eval-surveyA repository of papers and resources for evaluating large language models.1,450
nlpai-lab/kullmKorea University Large Language Model developed by researchers at Korea University and HIAI Research Institute.576
academic-hammer/hammerllmA large language model pre-trained on Chinese and English data, suitable for natural language processing tasks.43
safellama/plexiglassA toolkit to detect and protect against vulnerabilities in Large Language Models.122
damo-nlp-sg/m3examA benchmark for evaluating large language models in multiple languages and formats93
phodal/aigcDeveloping and applying large language models to improve software development workflows and processes1,413
csuhan/onellmA framework for training and fine-tuning multimodal language models on various data types601
trusted-ai/aix360A toolkit for explaining complex AI models and data-driven insights1,641
ucsc-vlaa/vllm-safety-benchmarkA benchmark for evaluating the safety and robustness of vision language models against adversarial attacks.72
michael-wzhu/shennong-tcm-llmDevelops and deploys a large language model for Chinese traditional medicine applications316
luogen1996/lavinAn open-source implementation of a vision-language instructed large language model513
aifeg/benchlmmAn open-source benchmarking framework for evaluating cross-style visual capability of large multimodal models84