📚Introduction to feature engineering for time series forecasting - extract from the book by Dr. Francesca “Machine Learning for Time Series Forecasting with Python” published by Wiley in December 2020
Developing ML models is often time-consuming and requires many factors to be considered: algorithm iteration, hyperparameter tuning, and feature engineering. These options are additionally multiplied by time series data, since DS specialists still need to take into account trends, seasonality, holidays, and external economic variables. Each ML algorithm expects data as input to be formatted. Therefore, time series datasets require precleaning and feature definitions before running the simulation.
The peculiarity of time series analysis is that data points are linked to time. For example, you need to build the output of an ML model by defining the variable you want to predict (future sales next Monday) and then use the historical data and feature set to create the input variables. This is necessary to achieve the following goals:
• creation of the correct set of input data for the ML-algorithm in order to create input features from historical data and form a dataset as a supervised learning problem;
• improving the performance of ML-models - creating a valid relationship between the input features and the target variable that needs to be predicted.
The main categories of features useful for time series analysis are as follows:
• Date time features based on the timestamp value of each observation, such as the hour, month, and day of the week for each observation. This also includes weekends and holidays, seasons, etc.
• Lag features and window features - Values at previous time steps that are considered useful because they are created on the assumption that what happened in the past may influence or contain some kind of internal information about the future. For example, it might be useful to create functions for sales that occurred on previous days at 4:00 pm if you want to predict similar sales at the same time the next day.
• Sliding (Rolling) Window Statistics, to compute statistics on values from a given sample of data by defining a range that includes the sample itself, as well as a specified number of values before and after the sample used.
• Expandable feature window statistics that include all previous data.
Illustrations and examples of Python are available here:
https://medium.com/data-science-at-microsoft/introduction-to-feature-engineering-for-time-series-forecasting-620aa55fcab0
Post #338
449