awesome-seml

Machine Learning guidelines

A curated list of articles and resources on software engineering best practices for building machine learning applications.

A curated list of articles that cover the software engineering best practices for building machine learning applications.

GitHub

1k stars
45 watching
123 forks
last commit: over 2 years ago
Linked from 3 awesome lists

awesomeawesome-listdeep-learningmachine-learningml-opssoftware-engineering

Awesome Software Engineering for Machine Learning / Broad Overviews

AI Engineering: 11 Foundational Practices⭐
Best Practices for Machine Learning Applications
Engineering Best Practices for Machine Learning⭐
Hidden Technical Debt in Machine Learning SystemsπŸŽ“β­
Rules of Machine Learning: Best Practices for ML Engineering⭐
Software Engineering for Machine Learning: A Case StudyπŸŽ“β­

Awesome Software Engineering for Machine Learning / Data Management

A Survey on Data Collection for Machine Learning A Big Data - AI Integration Perspective_2019πŸŽ“
Automating Large-Scale Data Quality VerificationπŸŽ“
Data management challenges in production machine learning
Data Validation for Machine LearningπŸŽ“
How to organize data labelling for ML
The curse of big data labeling and three ways to solve it
The Data Linter: Lightweight, Automated Sanity Checking for ML Data SetsπŸŽ“
The ultimate guide to data labeling for ML

Awesome Software Engineering for Machine Learning / Model Training

10 Best Practices for Deep Learning
Apples-to-apples in cross-validation studies: pitfalls in classifier performance measurementπŸŽ“
Fairness On The Ground: Applying Algorithmic FairnessApproaches To Production SystemsπŸŽ“
How do you manage your Machine Learning Experiments?
Machine Learning Testing: Survey, Landscapes and HorizonsπŸŽ“
Nitpicking Machine Learning Technical Debt
On Comparing Classifiers: Pitfalls to Avoid and a Recommended ApproachπŸŽ“β­
On human intellect and machine failures: Troubleshooting integrative machine learning systemsπŸŽ“
Pitfalls and Best Practices in Algorithm ConfigurationπŸŽ“
Pitfalls of supervised feature selectionπŸŽ“
Preparing and Architecting for Machine Learning
Preliminary Systematic Literature Review of Machine Learning System Development ProcessπŸŽ“
Software development best practices in a deep learning environment
Testing and Debugging in Machine Learning
What Went Wrong and Why? Diagnosing Situated Interaction Failures in the WildπŸŽ“

Awesome Software Engineering for Machine Learning / Deployment and Operation

Best Practices in Machine Learning Infrastructure
Building Continuous Integration Services for Machine LearningπŸŽ“
Continuous Delivery for Machine Learning⭐
Continuous Training for Production ML in the TensorFlow Extended (TFX) PlatformπŸŽ“
Fairness Indicators: Scalable Infrastructure for Fair ML SystemsπŸŽ“
Machine Learning Logistics
Machine learning: Moving from experiments to production
ML Ops: Machine Learning as an engineered disciplined
Model Governance Reducing the Anarchy of ProductionπŸŽ“
ModelOps: Cloud-based lifecycle management for reliable and trusted AI
Operational Machine Learning
Scaling Machine Learning as a ServiceπŸŽ“
TFX: A tensorflow-based Production-Scale ML PlatformπŸŽ“
The ML Test Score: A Rubric for ML Production Readiness and Technical Debt ReductionπŸŽ“
Underspecification Presents Challenges for Credibility in Modern Machine LearningπŸŽ“
Versioning for end-to-end machine learning pipelinesπŸŽ“

Awesome Software Engineering for Machine Learning / Social Aspects

Data Scientists in Software Teams: State of the Art and ChallengesπŸŽ“
Machine Learning Interviews9,227over 3 years ago
Managing Machine Learning Projects
Principled Machine Learning: Practices and Tools for Efficient Collaboration

Awesome Software Engineering for Machine Learning / Governance

A Human-Centered Interpretability Framework Based on Weight of EvidenceπŸŽ“
An Architectural Risk Analysis Of Machine Learning Systems
Beyond Debiasing
Closing the AI Accountability Gap: Defining an End-to-End Framework for Internal Algorithmic AuditingπŸŽ“
Inherent trade-offs in the fair determination of risk scoresπŸŽ“
Responsible AI practices⭐
Toward Trustworthy AI Development: Mechanisms for Supporting Verifiable Claims
Understanding Software-2.0πŸŽ“

Awesome Software Engineering for Machine Learning / Tooling

AimAim is an open source experiment tracking tool
AirflowProgrammatically author, schedule and monitor workflows
Alibi Detect2,262almost 2 years agoPython library focused on outlier, adversarial and drift detection
Archai468almost 2 years agoNeural architecture search
Data Version Control (DVC)DVC is a data and ML experiments management tool
Facets Overview / Facets DiveRobust visualizations to aid in understanding machine learning datasets
FairLearnA toolkit to assess and improve the fairness of machine learning models
Git Large File System (LFS)Replaces large files such as datasets with text pointers inside Git
Great Expectations10,054almost 2 years agoData validation and testing with integration in pipelines
HParams126almost 2 years agoA thoughtful approach to configuration management for machine learning projects
KubeflowA platform for data scientists who want to build and experiment with ML pipelines
Label Studio19,798almost 2 years agoA multi-type data labeling and annotation tool with standardized output format
LiFT167over 3 years agoLinkedin fairness toolkit
MLFlowManage the ML lifecycle, including experimentation, deployment, and a central model registry
Model Card Toolkit427about 3 years agoStreamlines and automates the generation of model cards; for model documentation
Neptune.aiExperiment tracking tool bringing organization and collaboration to data science projects
Neuraxle610over 3 years agoSklearn-like framework for hyperparameter tuning and AutoML in deep learning projects
OpenMLAn inclusive movement to build an open, organized, online ecosystem for machine learning
PyTorch Lightning28,636almost 2 years agoThe lightweight PyTorch wrapper for high-performance AI research. Scale your models, not the boilerplate
REVISE: REvealing VIsual biaSEs111about 4 years agoAutomatically detect bias in visual data sets
Robustness Metrics466about 2 years agoLightweight modules to evaluate the robustness of classification models
Seldon Core4,409almost 2 years agoAn MLOps framework to package, deploy, monitor and manage thousands of production machine learning models on Kubernetes
Spark Machine LearningSpark’s ML library consisting of common learning algorithms and utilities
TensorBoardTensorFlow's Visualization Toolkit
Tensorflow Extended (TFX)An end-to-end platform for deploying production ML pipelines
Tensorflow Data Validation (TFDV)766almost 2 years agoLibrary for exploring and validating machine learning data. Similar to Great Expectations, but for Tensorflow data
Weights & BiasesExperiment tracking, model optimization, and dataset versioning

Backlinks from these awesome lists:

More related projects: