Abstract:
Samoa (Scalable Advanced Massive Online Analysis) is a platform for mining big data streams. It provides a collection of distributed streaming algorithms for the most common data mining and machine learning tasks such as classification, clustering, and regression, as well as programming abstractions to develop new algorithms. It features a pluggablearchitecturethatallowsittorunonseveraldistributedstreamprocessingengines such as Storm, S4, and Samza. samoa is written in Java, is open source, and is available at http://samoa-project.net under the Apache Software License version 2.0. Keywords: data streams, distributed systems, classification, clustering, regression, toolbox, machine learning