Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Imputation

  • Missing data can occur in a series or stream in different ways:

    • In some cases, this means we need to have logic handling different events or types of data that are not always included, e.g., precipitation in weather data which could be rain or snow, possibly reported separately as with OpenWeatherMap.

    • For the rain/snow, the missing value is trivial, i.e., missing means 0.

    • Random dropouts of streams and randomly missing data due to sensor malfunctions, network glitches, maintenance, etc. can make subsequent analyses faulty.

    • If data are missing systematically, this can either make imputation easy (we know what the values should have been) or biased (the imputation is consistently wrong).

Simple imputation

  • For variables that we assume have a fixed distribution, imputing with the mean, trimmed mean or median value can be sufficient.

  • scikit-learn’s SimpleImputer contains the basic imputations.

length    1
dtype: int64
Loading...
length    0
dtype: int64
Loading...

Neighbour imputation

  • Instead of using “global” values, i.e., mean, median, etc., a different strategy is to take inspiration from neighbouring observations.

    • If a single variable is used, neighbours can be objects close in the sequence.

    • If multiple variables are used, neighbours can be objects with similar properties in the remaining (non-missing) variables.

  • The size of the neigbourhood, K, as in K Nearest Neighbours, can be used to smooth (large K) or ensure local adaption (small K).

  • scikit-learn’s KNNImputer works well on tabular data.

<Figure size 600x400 with 1 Axes>
[2.70866691]
[2.70866691]

Discussion point

  • What happened to our imputation above?

  • Why did it not work as expected?

Imputation in time series

  • Most of scikit-learn’s imputers are made for tabular data.

    • Each sample is seen as independent of the order.

    • Time series data are strictly ordered.

  • Imputation techniques for time series:

    • Last Observation Carried Forward (LOCF).

    • Next Observation Carried Backward (NOCB).

    • Interpolation, e.g., linear between neighbour points, local polynomial fitting, splines.

<Figure size 600x400 with 1 Axes>
The LOCF imputed value is -9.541817479586834

Interpolation

  • When a time series has consecutive dropouts, interpolation can make better imputations.

  • Pandas’ Series interpolation includes many different interpolations, e.g., linear, polynomial, splines, etc.

  • We will knock out a few points and compare some interpolations.

<Figure size 1200x400 with 2 Axes>
Loading...

Multivariate data imputation

  • For multivariate data it makes sense to leverage the other variables when one variable contains a missing value.

  • Nearest Neighbour imputation was mentioned above as a candidate.

  • An alternative is to use an iterative imputer, e.g., scikit-learns’ Iterative Imputer which predicts each variable from the other variables, imputing and remodelling iteratively.

Exercise

  • Use the DEWP, TEMP and PRES variables of the Beijing pollution data (2000 timepoints).

  • Remove timepoints 1000 to 1005 from the DEWP series.

  • Apply the iterative imputer to recreate the missing data.

  • Compare the results to a simple mean imputation by plotting the DEWP for a suitable region around the missing observations.

Notebook Cell