Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Streaming models

Static models

  • A pretrained model can be applied to a stream of data.

  • We will train a model using scikit-learn and predict in a loop.

  • Then try to convert it to a river streaming model before applying it to a stream of data.

  • Finally, create a streaming model directly in river.

Accuracy: 0.958 (± 0.035)
Loading...
Accuracy: 97.34%

Converting to river

Accuracy: 97.34%

Direct usage of river

  • The defaults in scikit-learn’s SGDClassifier and river’s LogisticRegression are different.

  • Pipelines can be made with a parenthesis and a pipe symbol.

Accuracy: 97.34%

Dynamic models

  • For comparison, we can continue learning during prediction of the test set (given that labels come together with the streaming data).

  • See the introduction on streaming models regarding strategies.

Accuracy: 96.81%

Inspect predictions one by one

<Figure size 1000x300 with 2 Axes>

Comment: In this case, one classification changed to the worse with the dynamic model, the rest stayed the same.

River vs scikit-learn

  • First of all, these are only examples of machine learning frameworks.

    • There are several other established and polished packages to use in their places, e.g., Quix, (scikit-multiflow + creme = River), PyTorch, TensorFlow, theano, PyCaret, OpenCV, etc.

  • River is built from the ground up with streaming data in mind.

    • Pre-processors, regressors and classifiers are all incremental.

    • A host of convenience functions for online/batch-wise learning are available.

  • scikit-learn is built for tabular data.

    • .partial_fit() is available for some pre-processors, regressors and classifiers.

    • Stream handling can be manually coded or helped by River and friends.

Live reading of Twitch chat, revisited

False

Synthetic streams

  • river can generate synthetic streams of various types.

Synthetic data generator Name Agrawal Task Binary classification Samples ∞ Features 9 Outputs 1 Classes 2 Sparse False Configuration ------------- classification_function 0 seed 42 balance_classes False perturbation 0.0
[103125.48379952488, 0, 21, 2, 8, 3, 319768.96426257194, 4, 338349.74371145567] 1
[135983.3438016299, 0, 25, 4, 14, 0, 423837.77555045625, 7, 116330.4466953698] 1
[98262.43477649744, 0, 55, 1, 18, 6, 144088.12440813935, 19, 139095.35411533137] 0
[133009.0417030814, 0, 68, 1, 14, 5, 233361.40250149256, 7, 478606.5361033906] 1
[63757.29086464148, 16955.938253511093, 26, 2, 12, 4, 522851.309309752, 24, 229712.43983592128] 1

Exercise

*HoeffdingTreeRegressor: This was originally called a Hoeffding Anytime Tree (HATT). It is an algorithm that is extremely efficient at updating decision trees with streaming data.

Streaming forecasts

  • river includes the SNARIMAX model, where N stands for non-linear, i.e., the (S)easonal (N)on-linear (A)uto(R)egressive (I)ntegrated (M)oving-(A)verage with e(X)ogenous inputs model.

  • The basic parameters match SARIMAX from the statsmodels package, but are named p/d/q/sp/sd/sq/m.

  • If no regressor is specified, a pipeline containing a StandardScaler and LinearRegression is used.

  • No statistics or summary tables are produced, so summaries must be manually created.

Airline passenger data

  • Monthly international passenger data from January 1949 through December 1960.

{'month': datetime.datetime(1949, 1, 1, 0, 0)} 112
{'month': datetime.datetime(1949, 2, 1, 0, 0)} 118
{'month': datetime.datetime(1949, 3, 1, 0, 0)} 132
{'month': datetime.datetime(1949, 4, 1, 0, 0)} 129
{'month': datetime.datetime(1949, 5, 1, 0, 0)} 121
{'month': datetime.datetime(1949, 6, 1, 0, 0)} 135
{'month': datetime.datetime(1949, 7, 1, 0, 0)} 148

Plot raw data

<Figure size 640x480 with 1 Axes>
1960-01-01 445.159 417.000
1960-02-01 476.332 391.000
1960-03-01 467.187 419.000
1960-04-01 452.182 461.000
1960-05-01 437.547 472.000
1960-06-01 438.079 535.000
1960-07-01 444.407 622.000
1960-08-01 452.582 606.000
1960-09-01 455.664 508.000
1960-10-01 455.428 461.000
1960-11-01 453.501 390.000
1960-12-01 452.975 432.000
R2: -0.254237
--

SARIMA + feature engineering

  • In addition to the original time series, we may add some freshly calculated exogenous variables.

  • In river’s SNARIMAX example, a distance function resembling a Radial Basis Function is applied to the months

    • This results in 12 new features measuring the distance to other months in the year.

    • In addition they include ordinal dates, i.e., day number since 0001-01-01.

{'January': 1.0, 'February': 0.36787944117144233, 'March': 0.01831563888873418, 'April': 0.00012340980408667956, 'May': 1.1253517471925912e-07, 'June': 1.3887943864964021e-11, 'July': 2.3195228302435696e-16, 'August': 5.242885663363464e-22, 'September': 1.603810890548638e-28, 'October': 6.639677199580735e-36, 'November': 3.720075976020836e-44, 'December': 2.820770088460135e-53, 'ordinal_date': 715510}
{'January': 0.00012340980408667956, 'February': 0.01831563888873418, 'March': 0.36787944117144233, 'April': 1.0, 'May': 0.36787944117144233, 'June': 0.01831563888873418, 'July': 0.00012340980408667956, 'August': 1.1253517471925912e-07, 'September': 1.3887943864964021e-11, 'October': 2.3195228302435696e-16, 'November': 5.242885663363464e-22, 'December': 1.603810890548638e-28, 'ordinal_date': 715601}
1960-01-01 418.156 417.000
1960-02-01 408.071 391.000
1960-03-01 441.154 419.000
1960-04-01 438.874 461.000
1960-05-01 448.431 472.000
1960-06-01 490.385 535.000
1960-07-01 532.368 622.000
1960-08-01 542.568 606.000
1960-09-01 488.254 508.000
1960-10-01 448.343 461.000
1960-11-01 414.920 390.000
1960-12-01 433.052 432.000
R2: 0.743524