Process mining is usually pitched at business processes, but the technique works on anything that produces a timestamped event log. This short piece looks at Google Borg trace data, the publicly released cluster workload traces, analysed in mindzie studio to identify bottlenecks and support capacity planning.

An interesting use case where we are leveraging the mindzie studio to process mine Google Borg trace data to identify bottlenecks and improve capacity allocation and scheduling decisions. The data provides information about 8 different Borg cells. It includes CPU usage, information about job to resource allocations, and job-parent information for MapReduce jobs. It is fully anonymized and does not contain any user information. The data can be obtained here:
https://github.com/google/cluster-data/blob/master/ClusterData2019.md
Arik Senderovich, PhD mindzie
Next steps
If your event data comes from infrastructure rather than an ERP, the approach is the same. Data Designer handles the log preparation and mindzie process mining handles the analysis. Try it with the free desktop edition.


