Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Outlier statistics

  • Determining if a sample or timepoint is a statistical outlier is often a two-step process:

    1. Estimate a distribution assumed to be normal operating conditions.

    2. Check if new samples are significant outliers from this distribution.

  • (Multivariate) Statistical Process Control, part of 6σ6\sigma process improvement, has a wide range of methods for this.

Statistical Process Control - SPC

  • The simplest form handles a single variable.

  • A normal distribution is assumed, but some deviation is tolerated.

  • It also assumes a constant/stationary process and no shifts/trends in distributions.

  • Any value outside +/- 3 standard deviations (SD) of the mean is assumed to be an outlier.

<Figure size 500x300 with 1 Axes>
Probability of being outside +/- 3 SD in a normal distribution: 0.27%
Proportion of values that are outside +/- 3 SD in a normal distribution: 1 in 370

Control charts

  • A common way of assessing outliers is to plot the series together with lower and upper boundaries.

<Figure size 1000x300 with 1 Axes>

Histograms

  • Histograms show you a cummulative view over a set of observations.

  • If the observed distribution is very different from a bell curve (normal distribution), the +/-3 SD will not have the expected coverage.

    • Associated P-values will be wrong.

<Figure size 500x300 with 1 Axes>
np.float64(7.538952188296564e-05)

Quantile plots

  • A quantile plot plots expected quantiles of a distributon against observed quantiles from data.

  • The basic distribution is the normal probability distribution.

<Figure size 400x300 with 1 Axes>

Robust statistics

  • Instead of mean and standard deviation, one can use more robust calculations.

  • Trimmed mean means first removing a proportion, e.g., 5% of the most extreme observations before calculating the mean.

    • Robust against outliers.

  • Median absolute deviation is the median of the absolute deviations from the median value.

    • Robust against non-normality.

MAD=median(∣Xi−X~∣),   X~=median(X)MAD = median(|X_i - \tilde{X}|), ~~~ \tilde{X} = median(X)
  • Relation to standard deviation: σ^=k⋅MAD\hat{\sigma} = k \cdot MAD.

  • For normal data k≈1.4826k \approx 1.4826, see Wikipedia.

Mean: 0.039
Trimmed mean: 0.046
Standard deviation: 0.981
Median absolute deviation: 0.647
Adjusted MAD: 0.959

Heuristics

  • As seen above, to be flagged in the base case will happen by accident 1 in 370 cases.

  • The probability quickly shrinks if additional requirements are added, e.g., 3 cases in a row outside +/- 3 std.

    • With multiple consecutive outliers, one can also use fewer standard deviations.

  • If data are assumed to be iid (independent and identically distributed), another heuristic can be to check if concecutive samples are too similar (and possibly non-centred).

    • This can indicate a bias in the series, e.g., caused by a manufacturing step caught in an error condition.

  • A shift in mean value can also be indicative of errors or unwanted changes; checkable using rolling means or similar with an appropriate window size.

    • A whole field of SPC uses Exponentially Weighted Moving Averages to detect outliers and shifts in distributions.

Probability: 0.0012%
Proportion: 1 in 84927
<Figure size 1000x300 with 1 Axes>

Exercise

  • Import the data called bananas.csv.

  • Assume the first 500 samples are “in control” and calculate their mean and standard deviation.

  • Plot the whole series and indicate outliers using SPC and heuristics.

Multivariate series

  • Though each variable can be handled separately with SPC, it is often more interesting to assess the combined effect.

  • Smaller deviations in single variables may be detected if these occur in relation to other variables and their variation.

  • Instead of the normal distribution, we use Mahalanobis distance and the associated Hotelling’s T2T^2, a multivariate generalisation of the student t-distribution. t2t^2 is defined as:

t2=(xˉ−μ)Σ^xˉ−1(xˉ−μ)Tt^2 = (\bar{x}-\mu) \hat{\Sigma}_{\bar{x}}^{-1} (\bar{x}-\mu)^T
  • In practice we estimate μ\mu and Σ^xˉ\hat{\Sigma}_{\bar{x}} from “in control” data and look at single observations for xˉ\bar{x}, i.e., xx.

n−pp(n−1)t2∼Fp,n−p\frac{n-p}{p(n-1)}t^2 \sim F_{p,n-p}

where nn is the number of samples in the training data and pp is the number of variables.

<Figure size 1000x600 with 3 Axes>

We use the first three years as the “in control” region and estimate mean and covariance from it.

[ -16.62009132   -5.42648402 1028.2803653 ]
[[ 33.50636014   8.27369045 -16.68247894]
 [  8.27369045  27.0510347  -12.9269373 ]
 [-16.68247894 -12.9269373   32.5255282 ]]
<Figure size 1000x600 with 2 Axes>

Finally, we observe which observations are marked as outlying in the original series.

<Figure size 1000x600 with 3 Axes>

Robust multivariate statistics

  • Also multivariate data can need robust statistics because of non-normality or outliers.

  • One method for robust estimation of mean and standard deviation is called Minimum Covariance Determinant (MCD).

  • MCD uses a subset of samples that minimises the determinant of the covariance matrix.

  • An iterative method that searches for the MCD is implemented in scikit-learn’s MinCovDet.

[ -14.55375772   -4.724472   1025.94736149]
[[ 42.75733002  21.71212165 -10.4415423 ]
 [ 21.71212165  44.64007986 -32.03879093]
 [-10.4415423  -32.03879093  25.19171101]]
[ -13.52995928   -2.41361257 1026.06748109]
[[ 50.39563984  16.364267   -33.07242179]
 [ 16.364267    31.89569218 -24.36020609]
 [-33.07242179 -24.36020609  48.7317869 ]]
[ -16.62009132   -5.42648402 1028.2803653 ]
[[ 33.50636014   8.27369045 -16.68247894]
 [  8.27369045  27.0510347  -12.9269373 ]
 [-16.68247894 -12.9269373   32.5255282 ]]

Exercise

  • Repeat the calculations of mean, covariance, Hotelling’s T2T^2, etc. for the Beijing pollution data.

  • Exchange ordinary statistics with the MCD alternative.

Time resolution

  • The (M)SPC examples have assumed a constant mean and standard deviation.

  • It may also be interesting to look at det deviation from smoothed data to look for local outliers.

    • This corresponds to a high-pass filter for FFT.

  • In the rich field of (M)SPC there are also other methods tailored for finding shifts in trends.

<Figure size 700x200 with 1 Axes>
<Figure size 700x400 with 2 Axes>