Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Machine Learning approach

  • Time series prediction, or forecasting, can be very similar to modelling and prediction with tabular data.

    • A set of input variables, usually a single output variable.

    • Ordinary machine learning methods can be applied.

  • The inclusion of time can be done by adding one or more delayed variables, possibly including the response.

Validation

  • As soon as time is part of a model, extra care needs to be taken in validation.

  • Cross-validation*, training-validation-test splits are still relevant.

  • However, training, validation and test sets need to follow time chronologically and cannot overlap.

  • Instead of traditional cross-validation, one can perform backtesting with a sliding or expanding window:

Figures from Roy Yang’s bloggpost on uber.com

Shipping, oil, interest rates and exchange rates

  • These data are public data from the Norwegian Bank, SSB, Eurostat and U.S. Energy Information Administration for the period 2000-2014 (monthly).

  • The data are available at ResearchGate and were part of a Master thesis by Raju Rimal.

/Users/kristian/Documents/GitHub/IND320/.venv/lib/python3.12/site-packages/openpyxl/worksheet/_reader.py:329: UserWarning: Unknown extension is not supported and will be removed
  warn(msg)
Loading...
Index(['Date', 'PerEURO', 'PerUSD', 'KeyIntRate', 'LoanIntRate', 'EuroIntRate', 'CPI', 'OilSpotPrice', 'ImpOldShip', 'ImpNewShip', 'ImpOilPlat', 'ImpExShipOilPlat', 'ExpCrdOil', 'ExpNatGas', 'ExpCond', 'ExpOldShip', 'ExpNewShip', 'ExpOilPlat', 'ExpExShipOilPlat', 'TrBal', 'TrBalExShipOilPlat', 'TrBalMland', 'ly.var', 'l2y.var', 'l.CPI', 'ExcChange', 'Testrain', 'season'], dtype='str')
/Users/kristian/Documents/GitHub/IND320/.venv/lib/python3.12/site-packages/openpyxl/worksheet/_reader.py:329: UserWarning: Unknown extension is not supported and will be removed
  warn(msg)
Loading...

Modelling without time

  • For starters, let us ignore time and build a simple prediction model for the exchange rate.

  • We will use scikit-learn’s Pipeline to combine standardisation (scaling) and linear regression and cross_val_predict to perform random K-fold cross-validation.

Loading...
<Figure size 640x480 with 1 Axes>
-0.9790095192601644
Cross-validated R2: -0.317

Backtesting

  • scikit-learn has a TimeSeriesSplit which creates segments for backtesting.

    • Expanding window is the default.

    • Sliding window can be applied by setting the right combination of parameters.

  • We will use scikit-learn’s cross_validate to perform the cross-validation based on the backtesting segments (cross_val_predict assumes that all observations will be test data at some point).

TimeSeriesSplit(gap=0, max_train_size=None, n_splits=5, test_size=None)
Fold 0:
  Train: index=[0]
  Test:  index=[1]
Fold 1:
  Train: index=[0 1]
  Test:  index=[2]
Fold 2:
  Train: index=[0 1 2]
  Test:  index=[3]
Fold 3:
  Train: index=[0 1 2 3]
  Test:  index=[4]
Fold 4:
  Train: index=[0 1 2 3 4]
  Test:  index=[5]
TimeSeriesSplit(gap=0, max_train_size=3, n_splits=3, test_size=None)
Fold 0:
  Train: index=[0 1 2]
  Test:  index=[3]
Fold 1:
  Train: index=[1 2 3]
  Test:  index=[4]
Fold 2:
  Train: index=[2 3 4]
  Test:  index=[5]
Fold 0:
  Train: index=[ 0  1  2  3  4  5  6  7  8  9 10 11 12 13 14 15]
  Test:  index=[16 17 18 19 20 21 22 23 24 25 26 27 28 29]
Fold 1:
  Train: index=[ 0  1  2  3  4  5  6  7  8  9 10 11 12 13 14 15 16 17 18 19 20 21 22 23
 24 25 26 27 28 29]
  Test:  index=[30 31 32 33 34 35 36 37 38 39 40 41 42 43]
Fold 2:
  Train: index=[ 0  1  2  3  4  5  6  7  8  9 10 11 12 13 14 15 16 17 18 19 20 21 22 23
 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43]
  Test:  index=[44 45 46 47 48 49 50 51 52 53 54 55 56 57]
Fold 3:
  Train: index=[ 0  1  2  3  4  5  6  7  8  9 10 11 12 13 14 15 16 17 18 19 20 21 22 23
 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47
 48 49 50 51 52 53 54 55 56 57]
  Test:  index=[58 59 60 61 62 63 64 65 66 67 68 69 70 71]
Fold 4:
  Train: index=[ 0  1  2  3  4  5  6  7  8  9 10 11 12 13 14 15 16 17 18 19 20 21 22 23
 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47
 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71]
  Test:  index=[72 73 74 75 76 77 78 79 80 81 82 83 84 85]
Fold 5:
  Train: index=[ 0  1  2  3  4  5  6  7  8  9 10 11 12 13 14 15 16 17 18 19 20 21 22 23
 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47
 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71
 72 73 74 75 76 77 78 79 80 81 82 83 84 85]
  Test:  index=[86 87 88 89 90 91 92 93 94 95 96 97 98 99]
Fold 6:
  Train: index=[ 0  1  2  3  4  5  6  7  8  9 10 11 12 13 14 15 16 17 18 19 20 21 22 23
 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47
 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71
 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95
 96 97 98 99]
  Test:  index=[100 101 102 103 104 105 106 107 108 109 110 111 112 113]
Fold 7:
  Train: index=[  0   1   2   3   4   5   6   7   8   9  10  11  12  13  14  15  16  17
  18  19  20  21  22  23  24  25  26  27  28  29  30  31  32  33  34  35
  36  37  38  39  40  41  42  43  44  45  46  47  48  49  50  51  52  53
  54  55  56  57  58  59  60  61  62  63  64  65  66  67  68  69  70  71
  72  73  74  75  76  77  78  79  80  81  82  83  84  85  86  87  88  89
  90  91  92  93  94  95  96  97  98  99 100 101 102 103 104 105 106 107
 108 109 110 111 112 113]
  Test:  index=[114 115 116 117 118 119 120 121 122 123 124 125 126 127]
Fold 8:
  Train: index=[  0   1   2   3   4   5   6   7   8   9  10  11  12  13  14  15  16  17
  18  19  20  21  22  23  24  25  26  27  28  29  30  31  32  33  34  35
  36  37  38  39  40  41  42  43  44  45  46  47  48  49  50  51  52  53
  54  55  56  57  58  59  60  61  62  63  64  65  66  67  68  69  70  71
  72  73  74  75  76  77  78  79  80  81  82  83  84  85  86  87  88  89
  90  91  92  93  94  95  96  97  98  99 100 101 102 103 104 105 106 107
 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125
 126 127]
  Test:  index=[128 129 130 131 132 133 134 135 136 137 138 139 140 141]
Fold 9:
  Train: index=[  0   1   2   3   4   5   6   7   8   9  10  11  12  13  14  15  16  17
  18  19  20  21  22  23  24  25  26  27  28  29  30  31  32  33  34  35
  36  37  38  39  40  41  42  43  44  45  46  47  48  49  50  51  52  53
  54  55  56  57  58  59  60  61  62  63  64  65  66  67  68  69  70  71
  72  73  74  75  76  77  78  79  80  81  82  83  84  85  86  87  88  89
  90  91  92  93  94  95  96  97  98  99 100 101 102 103 104 105 106 107
 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125
 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141]
  Test:  index=[142 143 144 145 146 147 148 149 150 151 152 153 154 155]
{'fit_time': array([0.0034461 , 0.00206327, 0.00166082, 0.0033648 , 0.00181198, 0.00175905, 0.00186896, 0.00169396, 0.00247335, 0.00195909]), 'score_time': array([0.00135779, 0.00083685, 0.00075603, 0.00183201, 0.00081897, 0.00078726, 0.00079513, 0.00079513, 0.00439095, 0.00083303]), 'test_score': array([-2.58120320e+03, -4.49066361e+00, -7.37304826e+00, 1.20667734e-01, -5.47564370e+00, -1.91343710e+01, -1.62104880e+00, -2.82911176e-02, -3.42287951e+00, -1.33079128e+01]), 'train_score': array([1. , 0.95668789, 0.92040214, 0.85950793, 0.84569018, 0.78860664, 0.69881999, 0.68809706, 0.65576638, 0.64771763])}
<Figure size 640x480 with 2 Axes>

Question: Does the behaviour make sense with regard to what is included in and predicted from the model?

Fold 0:
  Train: index=[ 0  1  2  3  4  5  6  7  8  9 10 11 12 13 14 15]
  Test:  index=[16 17 18 19 20 21 22 23 24 25 26 27 28 29]
Fold 1:
  Train: index=[ 0  1  2  3  4  5  6  7  8  9 10 11 12 13 14 15 16 17 18 19 20 21 22 23
 24 25 26 27 28 29]
  Test:  index=[30 31 32 33 34 35 36 37 38 39 40 41 42 43]
Fold 2:
  Train: index=[ 0  1  2  3  4  5  6  7  8  9 10 11 12 13 14 15 16 17 18 19 20 21 22 23
 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43]
  Test:  index=[44 45 46 47 48 49 50 51 52 53 54 55 56 57]
Fold 3:
  Train: index=[13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36
 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57]
  Test:  index=[58 59 60 61 62 63 64 65 66 67 68 69 70 71]
Fold 4:
  Train: index=[27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50
 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71]
  Test:  index=[72 73 74 75 76 77 78 79 80 81 82 83 84 85]
Fold 5:
  Train: index=[41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64
 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85]
  Test:  index=[86 87 88 89 90 91 92 93 94 95 96 97 98 99]
Fold 6:
  Train: index=[55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78
 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99]
  Test:  index=[100 101 102 103 104 105 106 107 108 109 110 111 112 113]
Fold 7:
  Train: index=[ 69  70  71  72  73  74  75  76  77  78  79  80  81  82  83  84  85  86
  87  88  89  90  91  92  93  94  95  96  97  98  99 100 101 102 103 104
 105 106 107 108 109 110 111 112 113]
  Test:  index=[114 115 116 117 118 119 120 121 122 123 124 125 126 127]
Fold 8:
  Train: index=[ 83  84  85  86  87  88  89  90  91  92  93  94  95  96  97  98  99 100
 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118
 119 120 121 122 123 124 125 126 127]
  Test:  index=[128 129 130 131 132 133 134 135 136 137 138 139 140 141]
Fold 9:
  Train: index=[ 97  98  99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114
 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132
 133 134 135 136 137 138 139 140 141]
  Test:  index=[142 143 144 145 146 147 148 149 150 151 152 153 154 155]
{'fit_time': array([0.01077986, 0.00419188, 0.00206995, 0.00275826, 0.00767589, 0.00313282, 0.00351477, 0.00156379, 0.00197911, 0.00196409]), 'score_time': array([0.00166607, 0.00104213, 0.00108099, 0.00119996, 0.00208116, 0.00100899, 0.00082731, 0.00070882, 0.00081372, 0.00159621]), 'test_score': array([-2.69048314e+03, -1.24210053e+00, -1.33007709e+00, 6.61087698e-01, -8.82686693e+00, -3.59907774e+02, 5.20015261e-01, -2.78711930e+00, -7.94540852e+04, 4.04120248e-01]), 'train_score': array([1. , 0.99034575, 0.95542869, 0.94013477, 0.95627895, 0.8961227 , 0.75351647, 0.96129276, 0.93446863, 0.94984244])}
<Figure size 640x480 with 2 Axes>

Question: Again; does the behaviour make sense with regard to what is included in and predicted from the model?

Exercise

  • Repeat the PerEuro predictions, but exchange LinearRegression with scikit-learns’s PLSRegression.

  • Check if the number of components in the PLS model has an effect on the explained variance (R2\text{R}^2), either manually or using a GridSearchCV.

Including the response variable in the predictors

  • As long as the training and test sets are not overlapping, we can include the response as a predictor.

  • Adding the response lagged can be done as a single variable or several variables (i.e., several different lags).

  • We will later look at ARIMA-type models where time lag is the main mechanism for modelling.

Loading...
{'fit_time': array([0.00292015, 0.00259471, 0.00361323, 0.00179982, 0.00196004, 0.00201821, 0.00193214, 0.00219798, 0.00187778, 0.00182295]), 'score_time': array([0.00140882, 0.00117397, 0.00091863, 0.00080109, 0.00082088, 0.00094676, 0.00081706, 0.00084591, 0.0011611 , 0.00093603]), 'test_score': array([-3.23917569e+00, -1.33159013e+00, -2.94524064e+00, 7.05745562e-01, -7.71426462e+00, -4.00119090e+02, 5.94148027e-01, -2.01189464e+00, -1.67957479e+04, 5.67210980e-01]), 'train_score': array([1. , 0.98737617, 0.94958633, 0.93515466, 0.95150022, 0.8785855 , 0.75049267, 0.95403986, 0.93250471, 0.94743916])}
<Figure size 640x480 with 2 Axes>

Five lags

Loading...
{'fit_time': array([0.00280404, 0.0025022 , 0.00299215, 0.0042491 , 0.00213289, 0.00206828, 0.00184917, 0.00195408, 0.00385094, 0.00223804]), 'score_time': array([0.00087714, 0.00094914, 0.00095701, 0.00236487, 0.00089097, 0.00086474, 0.00082564, 0.00087094, 0.00119877, 0.00092983]), 'test_score': array([-5.69820756e-01, -1.57221588e+00, -4.97861817e+00, 7.36545452e-02, -7.18201630e+00, -7.06614762e+02, 5.51939962e-01, -2.81878260e+00, -2.51922682e+00, 4.95828753e-01]), 'train_score': array([1. , 0.99248991, 0.96940659, 0.94583791, 0.95749656, 0.91409406, 0.76327123, 0.97259431, 0.95439249, 0.96453061])}
<Figure size 640x480 with 2 Axes>

Correlation between time series

  • To get an impression of the connection between different variables, one can compute correlations, e.g., in the form of a correlation matrix.

  • If one expects one variable to affect another variable at a later time, correlation with a lag can be computed.

  • The degree of connection between two time series may also be dependent on time.

    • A Sliding Window Correlation (SWC) shows local correlation in time windows.

    • The window size (and possible lag) can be tuned for series of quick or slow changes.

  • Note: Correlation does not equal causation.

    • There may not be a cause and effect, even though two phenomena show similar patterns. Beautifully illustrated by Tyler Vigen.

  • The concept of Autocorrelation will be covered later.

Correlation between PerEURO and ExpNatGas: -0.000
Correlation between PerEURO and ExpNatGas lagged 10 timepoints: 0.093
Correlation between ExpNatGas and PerEURO lagged 0 timepoints: -0.000
Loading...
<Figure size 640x480 with 3 Axes>
Loading...
<function __main__.plot_SWC(center=22)>

Pandas’ rolling() and shifts

  • When applying Pandas’ rolling() function, the index is used for matching the data points.

  • Therefore, we need to shift the index of the ExpNatGas to achieve a lag.

  • Because of the sliding window, the two series do not need to match in length.

<Figure size 640x480 with 1 Axes>
RangeIndex(start=10, stop=189, step=1)

Exercise

  • Combine lag and sliding window correlation.

  • Use ipywidgets to control:

    • window width

    • lag

    • selected variable to compare to PerEURO

    • bonus: visualize the sliding window like above