Process Mining

Process Mining Google Borg Trace Data

Process mining

Process mining is usually pitched at business processes, but the technique works on anything that produces a timestamped event log. This short piece looks at Google Borg trace data, the publicly released cluster workload traces, analysed in mindzie studio to identify bottlenecks and support capacity planning.

Process mining

An interesting use case where we are leveraging the mindzie studio to process mine Google Borg trace data to identify bottlenecks and improve capacity allocation and scheduling decisions. The data provides information about 8 different Borg cells. It includes CPU usage, information about job to resource allocations, and job-parent information for MapReduce jobs. It is fully anonymized and does not contain any user information. The data can be obtained here:

https://github.com/google/cluster-data/blob/master/ClusterData2019.md

Arik Senderovich, PhD mindzie

Next steps

If your event data comes from infrastructure rather than an ERP, the approach is the same. Data Designer handles the log preparation and mindzie process mining handles the analysis. Try it with the free desktop edition.

Related reading

About the Author

Archives

Recent Articles