<!-- llms-explorer concept facts · https://llms-explorer.com/tree/anomaly-detection/ · pack 2026-09-08 · ~6081 tokens -->

# anomaly detection

> The discipline of separating "normal" from "not normal" when you mostly only have examples of normal. This skill covers the working methods, when each fits, and the gotchas that bite teams in producti

Parent: [Data Analysis](https://llms-explorer.com/tree/data-analysis/) · 32 facets · 89 facts · page: https://llms-explorer.com/tree/anomaly-detection/

## Anomaly Detection

- The discipline of separating "normal" from "not normal" when you mostly only have examples of normal. This skill covers the working methods, when each fits, and the gotchas that bite teams in production. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#anomaly-detection)

## When to use this skill

- Activate when the user: — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#when-to-use-this-skill)
  - is looking for unusual rows / events / time points — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#when-to-use-this-skill)
  - is setting up monitoring with alerting on a metric stream — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#when-to-use-this-skill)
  - is building fraud / fault / intrusion detection — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#when-to-use-this-skill)
  - needs to compare methods (Isolation Forest vs LOF vs autoencoder) — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#when-to-use-this-skill)
  - needs streaming anomaly detection — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#when-to-use-this-skill)
  - needs to distinguish data drift from anomalies — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#when-to-use-this-skill)

## When NOT to use this skill

- Forecasting → da-analytical-methods (references/da-15-forecasting.md) — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#when-not-to-use-this-skill)
- Supervised classification on labeled fraud → da-analytical-methods (references/da-7-machine-learning.md) — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#when-not-to-use-this-skill)
- Outlier spot-check during cleaning → da-analytical-methods (references/da-4-data-cleaning-preparation.md or references/da-5-exploratory-data-analysis.md) — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#when-not-to-use-this-skill)
- Causal investigation → da-analytical-methods (references/da-12-ab-testing-causal-inference.md) — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#when-not-to-use-this-skill)

## Framing: three problem types

- Before picking a method, name the problem type. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#framing-three-problem-types)
- Methods don't transfer cleanly between types. A z-score finds point anomalies but misses contextual and collective ones. STL-residual analysis handles contextual time-series anomalies. Sequence models or windowed statistics handle collective. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#framing-three-problem-types)

## z-score

- z = (x - μ) / σ. Flag if |z| > 3. Assumes approximately normal; sensitive to the very outliers you're trying to find (μ and σ get pulled). — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#z-score)

## Modified z-score (MAD-based)

- z_mod = 0.6745 × (x - median) / MAD. Flag if |z_mod| > 3.5 (Iglewicz & Hoaglin 1993). Robust to outliers because median and MAD don't move much. Use this instead of plain z-score. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#modified-z-score-mad-based)

## Grubbs's test

- Tests whether the single most extreme point is an outlier under a normality assumption. Tests one at a time; for multiple outliers use ESD. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#grubbss-test)

## Generalized ESD (Rosner 1983)

- Iteratively tests up to k suspected outliers in a normal sample. Computes test statistic for the most extreme point, removes it, repeats. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#generalized-esd-rosner-1983)

## IQR / Tukey fences

- lower = Q1 - 1.5·IQR, upper = Q3 + 1.5·IQR. Used by boxplots. Robust to outliers, no distribution assumption, but not statistically calibrated. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#iqr-tukey-fences)
- When to reach for each: modified z-score for clean tabular numerical data, IQR for a quick exploratory boxplot, ESD for the formal "are there k outliers in this sample" answer, Grubbs only for the single-outlier case. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#iqr-tukey-fences)

## Control charts: CUSUM and EWMA

  - CUSUM (Cumulative Sum) - accumulates deviations from the target. Triggers when the cumulative sum exceeds a threshold. Best for small persistent shifts. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#control-charts-cusum-and-ewma)
  - EWMA (Exponentially Weighted Moving Average) - exponentially-weighted average crosses control limits. Smoother than CUSUM; good for medium drifts. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#control-charts-cusum-and-ewma)
  - Shewhart 3σ - the classic; sensitive to single large jumps but slow on small persistent shifts. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#control-charts-cusum-and-ewma)
- These come from manufacturing SPC (statistical process control) but transfer to any monitored stream. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#control-charts-cusum-and-ewma)

## Change-point detection

- When the distribution changes, not just one point. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#change-point-detection)
- Use change-point detection when "anomaly" really means "this segment is from a different distribution than the previous segment." — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#change-point-detection)

## STL residual analysis

- Decompose the series via STL (statsmodels.tsa.seasonal.STL) into trend + seasonality + residual. Apply a point-anomaly method to the residual. This automatically handles seasonality, so you don't false-alarm on every December spike. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#stl-residual-analysis)

## k-NN distance

- Distance to the k-th nearest neighbor. Big distance = anomaly. Simple, works in low dimensions, scales badly past ~50 features. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#k-nn-distance)

## LOF — Local Outlier Factor (Breunig 2000)

- A point's anomaly score is the ratio of its local density to the local density of its neighbors. Catches anomalies in non-uniform-density data where global thresholds fail. Implemented in scikit-learn. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#lof-local-outlier-factor-breunig-2000)

## DBSCAN as outlier detector

- Density-based clustering - anything not in a dense region is a "noise" point. Outlier detection is a free side-effect. Sensitive to eps and min_samples. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#dbscan-as-outlier-detector)

## Isolation Forest (Liu, Ting, Zhou 2008)

- Build random trees by randomly picking a feature and a random split until each point is isolated. Anomalies have shorter average path lengths because random splits separate them quickly. Linear time, constant memory, the default for tabular numerical data above a few features. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#isolation-forest-liu-ting-zhou-2008)
- Hyperparameters: n_estimators=100 (default fine), max_samples=256 (canonical), contamination (your guess at anomaly rate; affects threshold). — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#isolation-forest-liu-ting-zhou-2008)

## Extended Isolation Forest (Hariri et al 2019)

- Fixes a known IF flaw: standard IF only splits on axes, biasing it on rotated data. EIF allows arbitrary hyperplane splits. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#extended-isolation-forest-hariri-et-al-2019)

## One-Class SVM

- Fits a boundary that encloses most of the training data. Anomalies fall outside the boundary. Sensitive to the nu parameter and kernel choice. Slow on > ~10k samples. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#one-class-svm)

## Mahalanobis distance / elliptic envelope

- Assumes Gaussian distribution; fits a covariance matrix; distance from the center weighted by the inverse covariance. Works on roughly elliptical data. EllipticEnvelope in scikit-learn uses robust covariance estimation (MCD - Minimum Covariance Determinant) so it isn't pulled by the very outliers you're trying to find. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#mahalanobis-distance-elliptic-envelope)

## Autoencoder reconstruction error

- Train an autoencoder on normal data. At inference, reconstruction error = anomaly score. Works because the model never learned to reconstruct rare patterns. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#autoencoder-reconstruction-error)

## VAE (Variational Autoencoder)

- Same idea but with a probabilistic latent space. The likelihood of the data under the model is the anomaly score. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#vae-variational-autoencoder)

## GAN-based (AnoGAN, GANomaly, f-AnoGAN)

- Train a GAN on normal data. Anomaly score from the difference between the input and the closest sample the generator can produce. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#gan-based-anogan-ganomaly-f-anogan)

## Transformer-based and time-series foundation models

- 2024-2026 frontier. Models like Anomaly-Transformer, TranAD, and time-series foundation models (Chronos, Moirai, TimesFM) can be adapted for anomaly detection by computing prediction error or likelihood under the model. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#transformer-based-and-time-series-foundation-models)
- When deep learning is overkill: if your data is < 10 features and < 100k rows, Isolation Forest or LOF will outperform a neural net while running in seconds. Reach for deep methods when you have images, audio, dense time series with structure, or millions of features. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#transformer-based-and-time-series-foundation-models)

## Streaming and real-time

- In production you rarely batch-score; you score one event at a time. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#streaming-and-real-time)
- Production constraints: — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#streaming-and-real-time)
  - Memory - streaming detectors must bound state (e.g., reservoir sampling) — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#streaming-and-real-time)
  - Latency - score in microseconds for fraud, milliseconds for monitoring — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#streaming-and-real-time)
  - Concept drift - distribution shifts over time; the detector must adapt — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#streaming-and-real-time)

## Drift vs anomaly — the critical distinction

- These look similar but require different responses. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#drift-vs-anomaly-the-critical-distinction)
- Production ML systems need both. Most monitoring failures come from confusing them. For full drift tooling (PSI, KS, NannyML, Evidently), see da-analytical-methods (references/da-42-ml-model-monitoring.md). — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#drift-vs-anomaly-the-critical-distinction)

## Evaluating anomaly detectors

- The hard part: by definition, anomalies are rare, so you usually don't have labeled validation data. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#evaluating-anomaly-detectors)
- When you do have labels (post-hoc): use precision-recall, F1, PR-AUC. Accuracy is meaningless because the class is imbalanced. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#evaluating-anomaly-detectors)
- When you don't have labels: use known synthetic anomalies, or use the time-shifted holdout where you assume the holdout had a similar anomaly rate. Or measure proxy metrics like "% of incidents the system caught." — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#evaluating-anomaly-detectors)
- The threshold choice is usually the hardest decision. The model emits a score; you choose where to cut. Tune for the cost-benefit ratio: if a false positive costs 1 minute of investigation and a false negative costs $10k, the threshold should be aggressive. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#evaluating-anomaly-detectors)

## Anti-patterns

- Z-score on data full of outliers - μ and σ are dragged; use modified z (MAD). — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#anti-patterns)
- Single threshold on a seasonal series - false-alarms on every Monday or every December. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#anti-patterns)
- Confusing drift with anomaly - retraining on the anomaly, or investigating drift as if it were an event. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#anti-patterns)
- Autoencoder for 5-feature tabular - overkill; IF will outperform with seconds of compute. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#anti-patterns)
- No baseline period - declaring everything new "anomalous" when you simply lack history. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#anti-patterns)
- Treating anomaly score as a probability - most methods produce uncalibrated scores; pick a threshold from PR data, not "p > 0.05". — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#anti-patterns)
- Alert fatigue - a noisy detector trains the on-call to ignore it. Tune precision before deploying. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#anti-patterns)
- Forgetting concept drift - the model that worked last quarter no longer represents "normal." — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#anti-patterns)

## References

- Chandola, V., Banerjee, A., & Kumar, V. (2009). "Anomaly Detection: A Survey." ACM Computing Surveys. The canonical survey. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#references)
- Iglewicz, B. & Hoaglin, D. (1993). How to Detect and Handle Outliers. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#references)
- Rosner, B. (1983). "Percentage Points for a Generalized ESD Many-Outlier Procedure." Technometrics. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#references)
- Breunig, M. M. et al. (2000). "LOF: Identifying Density-Based Local Outliers." SIGMOD. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#references)
- Liu, F. T., Ting, K. M., & Zhou, Z.-H. (2008). "Isolation Forest." ICDM. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#references)
- Hariri, S., Carrasco Kind, M., Brunner, R. J. (2019). "Extended Isolation Forest." IEEE TKDE. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#references)
- Truong, C., Oudre, L., & Vayatis, N. (2020). "Selective review of offline change point detection methods." Signal Processing. (PELT survey.) — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#references)
- Adams, R. P. & MacKay, D. J. C. (2007). "Bayesian Online Changepoint Detection." arXiv:0710.3742. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#references)
- River (online ML) - https://riverml.xyz/ — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#references)
- PySAD - https://github.com/selimfirat/pysad — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#references)
- Xu, J. et al. (2021). "Anomaly Transformer." ICLR. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#references)
- Schölkopf, B. et al. (2001). "Estimating the Support of a High-Dimensional Distribution." (One-Class SVM.) — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#references)
- scikit-learn outlier detection - https://scikit-learn.org/stable/modules/outlier_detection.html — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#references)
- Evidently AI drift detection guide - https://docs.evidentlyai.com/ (drift-vs-anomaly framing). — [source](https://llms-explorer.com/sources/mdb-context-hub/da-16-anomaly-detection/#references)

## Where this helps

- Setting up alerting on a metric stream where you mostly only have examples of 'normal' and need to flag point, contextual, or collective deviations without labeled anomaly data. — [source](https://llms-explorer.com/tree/anomaly-detection/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Choosing between a fast statistical method (modified z-score, IQR, control charts) and a model-based method (Isolation Forest, LOF, an autoencoder) based on data size, dimensionality, and whether the data is seasonal. — [source](https://llms-explorer.com/tree/anomaly-detection/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Distinguishing a genuine anomaly from concept drift in a production ML system — the two look similar but call for opposite responses, and confusing them is this pack's most-cited monitoring failure mode. — [source](https://llms-explorer.com/tree/anomaly-detection/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Building fraud, fault, or intrusion detection where labeled positive examples are rare or nonexistent and the detector has to be evaluated without a clean validation set. — [source](https://llms-explorer.com/tree/anomaly-detection/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*

## Project ideas

- Build a seasonality-aware alerting pipeline: decompose the series with STL, apply modified z-score or IQR to the residual, and compare false-alarm rates against a naive single-threshold detector on the raw series. — [source](https://llms-explorer.com/tree/anomaly-detection/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Implement an Isolation Forest baseline for tabular fraud or fault detection (n_estimators=100, max_samples=256) and only reach for an autoencoder or deep method if the data has more structure (images, audio, dense time series) than IF can exploit. — [source](https://llms-explorer.com/tree/anomaly-detection/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Build a streaming anomaly detector with bounded memory (reservoir sampling) that scores one event at a time within a latency budget, then instrument it to separately flag concept drift versus point anomalies. — [source](https://llms-explorer.com/tree/anomaly-detection/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Set up a CUSUM or EWMA control chart on a monitored production metric to catch small persistent shifts that a Shewhart 3-sigma chart would miss. — [source](https://llms-explorer.com/tree/anomaly-detection/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*

## Common mistakes

- Running plain z-score on data that's already full of outliers, letting the mean and standard deviation get pulled by the very points you're trying to flag — modified z-score (MAD-based) is robust to exactly this. — [source](https://llms-explorer.com/tree/anomaly-detection/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Applying a single fixed threshold to a seasonal series, producing false alarms every Monday or every December instead of decomposing seasonality out first. — [source](https://llms-explorer.com/tree/anomaly-detection/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Reaching for an autoencoder or other deep method on a 5-feature tabular dataset, when Isolation Forest or LOF will outperform it while running in seconds. — [source](https://llms-explorer.com/tree/anomaly-detection/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Treating an anomaly score as a calibrated probability and thresholding at something like 'p > 0.05' — most methods produce uncalibrated scores, so the threshold should come from precision-recall data instead. — [source](https://llms-explorer.com/tree/anomaly-detection/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*

## Known issues

- By definition, anomalies are rare, so labeled validation data usually doesn't exist — evaluation typically falls back to synthetic anomalies, time-shifted holdouts, or proxy metrics like the percentage of known incidents caught. — [source](https://llms-explorer.com/tree/anomaly-detection/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Threshold choice is usually the hardest decision in the whole pipeline and has to be tuned against the real cost-benefit ratio of a false positive versus a false negative, not picked from a generic rule of thumb. — [source](https://llms-explorer.com/tree/anomaly-detection/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Confusing drift with anomaly is the most common production failure mode this pack documents — retraining on what's actually drift, or investigating drift as if it were a discrete event, both waste effort and miss the real issue. — [source](https://llms-explorer.com/tree/anomaly-detection/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- An undertuned detector causes alert fatigue that trains the on-call rotation to ignore it, so precision has to be tuned before deployment, not after the team has already learned to dismiss the alerts. — [source](https://llms-explorer.com/tree/anomaly-detection/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*

## Context files

- [anomaly detection](https://llms-explorer.com/downloads/sources/mdb-context-hub/da-16-anomaly-detection.md)
