LM_Memorization

Content Extractor

A tool to extract memorized content from large language models like GPT-2 by analyzing their training data

Training data extraction on GPT-2

GitHub

179 stars
7 watching
33 forks
Language: Python
last commit: over 3 years ago

Related projects:

RepositoryDescriptionStars
ftramer/steal-mlA tool for extracting machine learning models from cloud-based services using prediction APIs344
eyurtsev/korAn open-source wrapper around LLMs to extract structured data from text1,638
recrm/archivetoolsA collection of tools for extracting and analyzing data from web archives71
iamgroot42/mimirA Python package for measuring memorization in Large Language Models.126
ir193/amextractorA tool to extract physical memory from Android devices without kernel source code or LKM support.12
kost/memdumpA tool to extract and display the contents of a system's physical memory12
eset-la/lord-of-the-stringsA tool to extract and classify relevant strings from binary files9
cognesy/instructor-phpA PHP library that simplifies the integration of Large Language Models into applications by providing structured data extraction and validation.230
halpomeranz/lmgTools and scripts for capturing and analyzing Linux memory266
gamallo/galextraA multi-language term extractor that uses morphosyntax tagging and filtering to identify multi-word terms from plain text input.2
bfelbo/deepmojiA deep learning model for analyzing sentiment and emotion in text based on emojis.1,525
os6sense/defmemoA macro that memoizes the results of functions with identical signatures33
monarch-initiative/ontogptAn LLM-based tool for extracting structured information from text with ontology-based grounding.626
knowledgecaptureanddiscovery/somefA tool that automatically extracts relevant metadata from code repositories, including software descriptions and bibliographic citations.47
yomurb/yomuA Ruby library for extracting text and metadata from various file formats.498