Skip to content
Appvizer
Nanny ML logo

Nanny ML : Post-deployment monitoring for ML model performance

Nanny ML: in summary

NannyML is an open-source Python library designed for post-deployment monitoring of machine learning models, specifically in scenarios where ground truth labels are delayed or unavailable. It is built for data scientists, ML engineers, and MLOps practitioners who need to assess model performance, detect data drift, and identify silent model failures in production.

Unlike traditional monitoring tools that rely on known target values, NannyML can estimate performance metrics even without real-time labels, using advanced statistical techniques. This makes it particularly valuable for applications such as credit scoring, fraud detection, and recommendation systems, where labels often arrive days or weeks after predictions are made.

Key benefits:

  • Estimates performance metrics without ground truth (e.g., estimated accuracy, precision, recall).
  • Detects data drift and feature importance changes.
  • Includes visual diagnostics and integrates with production ML workflows.

What are the main features of NannyML?

Estimation of model performance without labels

NannyML can monitor how well a model is performing in real-time, even before actual outcomes are known:

  • Uses Confidence-Based Performance Estimation (CBPE) and Direct Loss Estimation (DLE)
  • Estimates classification and regression metrics over time
  • Flags sudden drops in performance that would otherwise go unnoticed
  • Useful for high-latency feedback environments

Data drift detection

Tracks whether input data distributions have changed over time:

  • Monitors drift at feature-level and dataset-level
  • Supports drift metrics such as Jensen-Shannon divergence, PSI, and Wasserstein distance
  • Highlights which features contribute most to drift
  • Helps assess whether retraining is needed

Target distribution and realized performance tracking

Once labels become available, NannyML compares actual performance metrics against estimated ones:

  • Aligns realized vs. estimated performance curves
  • Evaluates calibration of estimation methods
  • Identifies discrepancies to refine monitoring strategies

Feature importance and data quality analysis

Provides insights into how and why model behavior changes:

  • Measures shifts in feature importance over time
  • Highlights missing or corrupted data in production
  • Assists in pinpointing data issues that impact model output

Report generation and visualization

NannyML produces interactive visual reports to support debugging and review:

  • Can be embedded in Jupyter notebooks or exported to HTML
  • Offers dashboards for temporal analysis and monitoring
  • Designed to explain anomalies clearly to technical teams

Why choose NannyML?

  • Performance monitoring without labels: Essential for use cases with delayed or unavailable outcomes.
  • Advanced statistical methods: Provides estimation techniques not commonly found in standard monitoring tools.
  • Open-source and framework-agnostic: Compatible with any model type or serving infrastructure.
  • Insightful diagnostics: Visual tools help interpret model behavior, drift, and failure causes.
  • Optimized for real-world production: Built to handle the challenges of monitoring ML in business-critical systems.

Nanny ML: its rates

Standard

Rate

On demand