trulens

Performance tracker for LLMs

A tool to evaluate and track the performance of large language model (LLM) experiments

Evaluation and Tracking for LLM Experiments

GitHub

2k stars
19 watching
197 forks
Language: Python
last commit: almost 2 years ago
Linked from 1 awesome list

explainable-mlllmllmopsmachine-learningneural-networks

Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
rlworkgroup/dowelA tool for logging and tracking machine learning research progress in Python32
replicate/keepsakeA tool for version control of machine learning experiments1,649
qcri/llmebenchA benchmarking framework for large language models81
polyaxon/tracemlA software framework for tracking and analyzing machine learning data, including model performance, inputs, outputs, and project metadata.510
didactic-drunk/fiber_metrics.crA tool to track and measure runtime, memory allocation, and other performance metrics for individual fibers or methods within concurrent applications.8
neptune-ai/neptune-clientAn experiment tracker for machine learning model training that allows users to log and visualize their experiments in detail.590
damo-nlp-sg/m3examA benchmark for evaluating large language models in multiple languages and formats93
catalyst-team/alchemyProvides tools and infrastructure to log and visualize experiments in deep learning research50
iamgroot42/mimirA Python package for measuring memorization in Large Language Models.126
freedomintelligence/mllm-benchEvaluates and compares the performance of multimodal large language models on various tasks56
alephbet/gimelAn A/B testing backend built using AWS Lambda and Redis HyperLogLog to efficiently track experiment data in a scalable and cost-effective manner.227
rwdaigle/metrixAn Elixir library to log custom application metrics in a well-structured format for downstream processing systems.52
trubrics/trubrics-sdkAn analytics platform designed to track performance and feedback for AI-powered assistants using machine learning and data analysis techniques.136
leks-forever/nllb-tuningThis is an experimental project for fine-tuning the NLB language model with a specific dataset and evaluating its performance on translation tasks.7
luogen1996/lavinAn open-source implementation of a vision-language instructed large language model513