🐱Visualizing Temporal Changes of Categorical Data: PyCatFlow vs RankFlow
Sometimes a Data Scientist needs to visualize ranked lists over time, such as changes in search results for queries on Google or YouTube. To do this, you can use RankFlow - a useful tool with a minimalistic UI and a rather cumbersome data preparation process. RankFlow allows you to compare ranked listings over time. It requires that the input tabular data be organized so that each column represents a ranked list. Each ranked list can be supplemented with weights, adding another level of information to the data. For example, for YouTube search results, you can take views, upvotes, or upvotes ratio. Each column in a data table is represented as a stack of nodes, ordered according to rank in a given dataset. In addition, identical nodes are connected between columns. This highlights data continuity and changes, allowing pattern analysis.
Building a RankFlow visualization from this data requires modifying the dataset. For each version of the API, there must be a column containing a ranked list of permissions that are not ordered by any relevance metric. Therefore, ordering a RankFlow chart is a design decision, meaning items can be sorted alphabetically, by frequency in the dataset, or based on additional data.
In practice, adapting the data to the required RankFlow data structure is quite tedious. To speed up the pre- and post-processing of charts, you can write your own Python script that processes the XML data in the SVG file generated by RankFlow. An alternative is PyCatFlow, a visualization tool similar to RankFlow that works well for temporal data without explicit ranking information, but with potential additional categorical data. PyCatFlow is an open-source Python package that can be downloaded freely from Github.
https://medium.com/@bumatic/pycatflow-visualizing-categorical-data-over-time-b344102bcce2
https://github.com/bumatic/PyCatFlow
Post #352
466