📝Analyzing Time Series Data: 5 Tips for a Data Scientist
One of the most common mistakes beginners make in analyzing time series data is the assumption that the data has regular points and does not contain gaps. In practice, this is usually not confirmed and leads to incorrect results. In real datasets, data points are often missing, and the available ones are located unevenly or inconsistently. Therefore, before analyzing time series data, a preliminary preparation stage should be carried out:
• Understand the time range and detail of the time series by data points using dataset visualization;
• Compare the actual number of ticks in each time series with the number of expected ticks depending on the interval between points and the total length of the time series. This ratio is sometimes referred to as the duty cycle, which is the difference between the maximum and minimum timestamp divided by the point spacing. If this value is much less than 1, then a lot of data is missing.
• Filter out batches with low duty cycle by setting a limit, for example, 40% or whatever is appropriate for a specific task.
• Standardize the spacing between time series cues by upsampling to finer resolution.
• Fill upsampled gaps using an appropriate interpolation method such as last known value or linear / quadratic interpolation. In Apache Spark, you can use the applyInPandas method in the PySpark grouped dataframe for this, under the hood of which is pandasUDF, the performance of which is much higher than simple UDF functions due to more efficient data transfer through Apache Arrow and calculations through Pandas vectorization.
https://towardsdatascience.com/a-common-mistake-to-avoid-when-working-with-time-series-data-eedf60a8b4c1
Post #333
466