1 of 6

Crunchy Clusters

Jason Rigby, Baichuan Sun, Grischa Meyer

2 of 6

Objective

Data on the web does not have standardised access methods

Data needs to be integrated into tools

Tools can be made available on the cloud

Centralising access and processing allows simple cross-database analysis

We want to provide a cloud based data science and machine learning tool with simplified open data access methods provided

3 of 6

Analysis in the cloud

Spark is a powerful framework for analysing and learning from large datasets

There is a iPython notebook style interface for Spark called Zeppelin

Supports Scala, R, Python, and others out of the box

Google Docs style instant collaboration

Backed by Mesos based cluster to perform computations

4 of 6

Proposed system architecture

NeCTAR Openstack or any other Cloud

Mesos cluster - spanning any number of resources

Zeppelin 1

Zeppelin 2

Zeppelin x

Spark process user x

DIT4C fork

Compute instance

Compute instance

Compute instance

Compute instance

5 of 6

Data access goals and issues

“Open Data” suggests similar ease of access to Open Source, but reality is different

A public REST API is nice, an API following a spec, e.g. Swagger, is better

In our project we wanted to show a possible approach to making data available inside an analysis framework, ready to be analysed

6 of 6

Thank you

Our thanks to organisers and sponsors