bleurt

NLG evaluation metric

An evaluation metric for Natural Language Generation based on transfer learning.

BLEURT is a metric for Natural Language Generation based on transfer learning.

GitHub

705 stars
13 watching
85 forks
Language: Python
last commit: about 3 years ago
Linked from 1 awesome list


Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
maluuba/nlg-evalA toolset for evaluating and comparing natural language generation models1,350
benhamner/metricsProvides implementations of various supervised machine learning evaluation metrics in multiple programming languages.1,632
mlgroupjlu/llm-eval-surveyA repository of papers and resources for evaluating large language models.1,450
bllip/bllip-parserA statistical natural language parser used to generate grammatically correct sentences from unstructured text input.227
microsoft/prophetnetA collection of research implementations and models for natural language generation694
i-gallegos/fair-llm-benchmarkCompiles bias evaluation datasets and provides access to original data sources for large language models115
thiagocf05/webnlgProvides intermediate representations of data for NLG tasks like Discourse Ordering and Lexicalization69
nlgranger/seqtoolsA Python library to manipulate and transform indexable data49
dluebke/bpelstatsA tool for calculating and analyzing BPEL metrics0
simplenlg/simplenlgA Java API for generating natural language texts from syntactic forms810
intellabs/fastragA framework for efficient and optimized retrieval augmented generative pipelines using state-of-the-art LLMs and Information Retrieval.1,392
lartpang/pysodmetricsA library providing an implementation of various metrics for object segmentation and saliency detection in computer vision.150
opennlg/openbaA pre-trained language model designed for various NLP tasks, including dialogue generation, code completion, and retrieval.94
google-research/deep_opeProvides benchmarking policies and datasets for offline reinforcement learning85