Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Synchronisation

  • When merging time-dependent data from different sources, matching them well is important, but also comes with some choices.

  • If time series are to be used in Machine Learning, exact synchronisation is usually needed, i.e., equal timestamps on each variable.

  • For plotting purposes different time resolution, e.g., weeks vs months, may not be a problem as long as the different cycles match up.

<Figure size 700x200 with 1 Axes>

Accumulation

  • If a higher resolution time series is to be synchronized with a lower resolution series, some type of accumulation needs to be done.

  • Things to keep in mind:

    • Which type of accumulation should be done, e.g., sum for count data, mean for intensity data?

    • Should we use simple statistics, robust statistics or smoothed series data?

    • Are the low resolution recordings a series of snapshots (a) or an accumulation between (before (b)/after (d)) or around (c) timepoints (see illustration)?

      • This coresponds to the different uses of filters in the Noise reduction section.

      • ... which also means that it is possible to echange simple averages with other smoothers.

(72,)
array([30.97322385, 34.35672352, 39.15225169, 42.54404014, 36.29402279])

Question: Is the mean calculation above an accumulation of type a, b, c or d (as compared to the illustration)?

Loading...
days 2021-01-01 30.973224 2021-01-02 34.356724 2021-01-03 39.152252 2021-01-04 42.544040 2021-01-05 36.294023 Name: electricity, dtype: float64
Loading...
<Figure size 640x480 with 1 Axes>

Interpolation

  • Timepoints may not match as easily as with days and hours above.

  • If one series is shifted slightly, the series are irregular in timesteps or have non-intuitive intervals, interpolation is an alternative.

  • When interpolating an irregular sequence in Pandas, one may need to resample to a higher frequency, fill the missing values and resample to the final frequency (see example below).

              values
2000-01-04  0.490463
2000-01-07  1.274086
2000-01-08 -0.485563
2000-01-19  0.886186
2000-01-21 -0.331479
2000-01-28 -0.039743
2000-01-30  0.477419
2000-02-19  0.814175
2000-02-26 -0.835243
2000-03-09 -1.014372
2000-03-22  0.125960
2000-04-07 -0.459632
2000-04-12  0.412295
2000-04-17  0.722275
2000-04-24  0.476232
2000-05-13  0.772985
2000-05-17 -0.650375
2000-05-22 -1.162162
2000-05-30 -0.086559
2000-06-03  0.286860
2000-06-21  0.059377
2000-06-22 -0.928979
2000-06-27 -1.953670
2000-07-03  0.294007
2000-07-14 -1.251394
2000-07-17 -1.134486
2000-07-24  1.254263
2000-07-26  1.509980
2000-07-29 -0.198877
2000-08-04 -0.532344
2000-08-07 -0.045698
2000-08-26 -0.090330
2000-09-03  2.802757
2000-09-07 -3.711439
2000-09-08  0.041011
2000-09-09 -0.014421
2000-09-12 -1.781971
2000-09-16  0.336200
2000-09-28 -0.027597
2000-10-06  0.013211
2000-10-12  2.575837
2000-10-13 -1.140271
2000-10-23 -1.149615
2000-10-28  1.270556
2000-10-29  0.021922
2000-11-11  0.701834
2000-11-20  0.209048
2000-11-23 -0.972874
2000-11-30 -0.407151
2000-12-28 -0.473028
Loading...
<Figure size 700x200 with 1 Axes>
Loading...
<Figure size 700x200 with 1 Axes>

Time delays

  • Industrial processes often have a continuous or batchwise handling of raw materials into other materials or products.

    • When sensors record data along the production line, matching a piece of raw material to sensor readings can be done by adding a delay to the timestamp of the measurements early in the process or subtracting time from the later measurements.

  • For some processes, the delay is a known, fixed quantity.

    • For others the delay may be dependent on dynamic factors like raw material properties or process settings that add uncertainty to the time delay.

    • Synchronising such data, may require optimising correlations between sensors or using more advanced warping techniques.