Crunchy Clusters
Jason Rigby, Baichuan Sun, Grischa Meyer
Objective
Data on the web does not have standardised access methods
Data needs to be integrated into tools
Tools can be made available on the cloud
Centralising access and processing allows simple cross-database analysis
We want to provide a cloud based data science and machine learning tool with simplified open data access methods provided
Analysis in the cloud
Spark is a powerful framework for analysing and learning from large datasets
There is a iPython notebook style interface for Spark called Zeppelin
Supports Scala, R, Python, and others out of the box
Google Docs style instant collaboration
Backed by Mesos based cluster to perform computations
Proposed system architecture
NeCTAR Openstack or any other Cloud
Mesos cluster - spanning any number of resources
Zeppelin 1
Zeppelin 2
Zeppelin x
Spark process user x
DIT4C fork
Compute instance
Compute instance
Compute instance
Compute instance
Data access goals and issues
“Open Data” suggests similar ease of access to Open Source, but reality is different
A public REST API is nice, an API following a spec, e.g. Swagger, is better
In our project we wanted to show a possible approach to making data available inside an analysis framework, ready to be analysed
Thank you
Our thanks to organisers and sponsors