Skip to content
Machine Learning SalonFree learning resources

Infrastructure

Big data and cloud computing

Distributed storage, processing and tools for data at scale.

BerkeleyISchool Videos

The School of Information is both UC Berkeley's newest and its smallest school. Located in the center of campus, the I School is a graduate education and research community that crosses the boundaries of law, economics, sociology, business, library science, engineering, design, publishing, linguistics, computer science, and information science.

In UC Berkeley's historic South Hall, roughly 100 graduate students and 13 faculty members form a small, multidisciplinary collective of scholars and practitioners.

The I School offers a professional master's degree and an academic doctoral degree. Our master's program trains students for careers as information professionals and emphasizes small classes and project based learning. Our Ph.D. program equips scholars to develop solutions and shape policies that influence how people seek, use, and share information.

Playlists

12 VIDEOS

DataEDGE 2014

3 months ago

6 VIDEOS

Privacy & Big Data Workshop: Values and Governance

7 months ago

6 VIDEOS

Admissions FAQ

1 year ago

10 VIDEOS

DataEDGE Conference 2013

1 year ago

15 VIDEOS

Analyzing Big Data with Twitter (Info 290, Fall 2012)

2 years ago

10 VIDEOS

Social Data Revolution (Info 290A, Fall 2012)

2 years ago

11 VIDEOS

DataEDGE Conference 2012

2 years ago

www.youtube.com

Links before 24 oct 2014

Coursera: Mining Massive Datasets

We introduce the student to modern distributed file systems and MapReduce, including what distinguishes good MapReduce algorithms from good algorithms in general. The rest of the course is devoted to algorithms for extracting models and information from large datasets. Students will learn how Google's PageRank algorithm models importance of Web pages and some of the many extensions that have been used for a variety of purposes. We'll cover locality sensitive hashing, a bit of magic that allows you to find similar items in a set of items so large you cannot possibly compare each pair. When data is stored as a very large, sparse matrix, dimensionality reduction is often a good way to model the data, but standard approaches do not scale well; we'll talk about efficient approaches. Many other large scale algorithms are covered as well, as outlined in the course syllabus.

Course Syllabus

Week 1:

MapReduce

Link Analysis PageRank

Week 2:

Locality Sensitive Hashing Basics + Applications

Distance Measures

Nearest Neighbors

Frequent Itemsets

Week 3:

Data Stream Mining

Analysis of Large Graphs

Week 4:

Recommender Systems

Dimensionality Reduction

Week 5:

Clustering

Computational Advertising

Week 6:

support vector Machines

Decision Trees

MapReduce Algorithms

Week 7:

More About Link Analysis Topic specific PageRank, Link Spam.

More About Locality Sensitive Hashing

www.coursera.org

The Caltech JPL Summer School on Big Data Analytics

The anticipated schedule of lectures (subject to changes):

Each bullet bellow corresponds to a set of materials that includes approximately 2 hours of video lectures, various links and supplementary materials, plus some online, hands on exercises.

1. Introduction to the school. Software architectures. Introduction to Machine Learning.

2. Best programming practices. Information retrieval.

3. Introduction to R. Markov Chain Monte Carlo.

4. Statistical resampling and inference.

5. Databases.

6. Data visualization.

7. Clustering and classification.

8. Decision trees and random forests.

9. Dimensionality reduction. Closing remarks.

www.coursera.org

Apache Spark Machine Learning Library

MLlib is a Spark implementation of some common machine learning (ML) functionality, as well associated tests and data generators. MLlib currently supports four common types of machine learning problem settings, namely, binary classification, regression, clustering and collaborative filtering, as well as an underlying gradient descent optimization primitive.

spark.apache.org

Apache Spark Summit Videos, 2014

www.youtube.com

Apache Spark Summit exercises, 2013

Welcome to the Spark Summit hands on exercises. These exercises are adapted from similar exercises that were prepared for and run at AMP Camp Big Data Bootcamps. They were written by volunteer graduate students and postdocs in the UC Berkeley AMPLab. Many of those same graduate students are also volunteers here on the Spark Summit Training day team as well. The exercises we cover today will have you working directly with the Spark specific components of the AMPLab’s open source software stack, called the Berkeley Data Analytics Stack (BDAS).

spark summit.org

Apache Spark Summit Training, 2014

Course Prerequisites:

Laptop with WiFi capabilities

Java 6 or 7

TRACK A: Introduction to Apache Spark Workshop

INTRO EXERCISES

The Introduction to Apache Spark workshop is for users to learn the core Spark APIs. This session features hands on technical exercises to get developers up to speed in using Spark for data exploration, analysis, and building big data applications.

The integrated lecture and lab format covers the following topics:

Overview of Big Data and Spark

Installing Spark Locally

Using Spark’s Core APIs in Scala, Java, & Python

Building Spark Applications

Deploying on a Big Data Cluster

Building Applications for Multiple Platforms

TRACK B:Advanced Apache Spark Workshop

ADVANCED EXERCISES

The Advanced Apache Spark Workshop will cover advanced topics on architecture, tuning, and each of Spark’s high level libraries (including the latest features). Attendees will have the opportunity after the lunch break to work through labs on each of the libraries.

Some familiarity with Spark or MapReduce is expected, as this workshop will not cover basic Spark programming.

Topics covered include:

Advanced Spark Internals and Tuning, Reynold Xin, SLIDES

Spark SQL, Michael Armburst, SLIDES

Spark Streaming, Tathagata Das, SLIDES

MLlib, Ameet Talwalkar, SLIDES

GraphX, Ankur Dave, SLIDES

spark summit.org

Apache Mahout ML library

The Apache Mahout™ project's goal is to build a scalable machine learning library.

Currently Mahout supports mainly three use cases: Recommendation mining takes users' behavior and from that tries to find items users might like. Clustering takes e.g. text documents and groups them into groups of topically related documents. Classification learns from exisiting categorized documents what documents of a specific category look like and is able to assign unlabelled documents to the (hopefully) correct category.

mahout.apache.org

Apache Mahout on Javaworld

Enjoy machine learning with Mahout on Hadoop, 2014

Mahout brings the power of scalable processing to Hadoop's huge data sets

www.javaworld.com

Know this right now about Hadoop, 2014

From core elements like HDFS and YARN to ancillary tools like Zookeeper, Flume, and Sqoop, here's your cheat sheet and cartography of the ever expanding Hadoop ecosystem.

www.javaworld.com

MapReduce programming with Apache Hadoop, 2008

Process massive data sets in parallel on large clusters

www.javaworld.com

Deeplearning4j

Deeplearning4j is the first commercial grade deep learning library written in Java. It is meant to be used in business environments, rather than as a research tool for extensive data exploration. Deeplearning4j is most helpful in solving distinct problems, like identifying faces, voices, spam or e commerce fraud.

Deeplearning4j aims to be cutting edge plug and play, more convention than configuration. By following its conventions, you get an infinitely scalable deep learning architecture. The framework has a domain specific language (DSL) for neural networks, to turn their multiple knobs.

Deeplearning4j includes a distributed deep learning framework and a normal deep learning framework; i.e. it runs on a single thread as well. Training takes place in the cluster, which means it can process massive amounts of data. Nets are trained in parallel via iterative reduce.

The distributed framework is made for data input and neural net training at scale, and its output should be highly accurate predictive models.

By following the links at the bottom of each page, you will learn to set up, and train with sample data, several types of deep learning networks. These include single and multithread networks, Restricted Boltzmann machines, deep belief networks and Stacked Denoising Autoencoders.

For a quick introduction to neural nets, please see our overview.

deeplearning4j.org

Udacity opencourseware "Intro to Hadoop and MapReduce"

Course Summary

The Apache™ Hadoop® project develops open source software for reliable, scalable, distributed computing. Learn the fundamental principles behind it, and how you can use its power to make sense of your Big Data.

Why Take This Course?

• How Hadoop fits into the world (recognize the problems it solves)

• Understand the concepts of HDFS and MapReduce (find out how it solves the problems)

• Write MapReduce programs (see how we solve the problems)

• Practice solving problems on your own

www.udacity.com

Storm Apache

Apache Storm is a free and open source distributed realtime computation system. Storm makes it easy to reliably process unbounded streams of data, doing for realtime processing what Hadoop did for batch processing. Storm is simple, can be used with any programming language, and is a lot of fun to use!

storm.incubator.apache.org

storm.incubator.apache.org

Michael Viogiatzis Blog

How to spot first stories on Twitter using Storm

As a first blog post, I decided to describe a way to detect first stories (a.k.a new events) on Twitter as they happen. This work is part of the Thesis I wrote last year for my MSc in Computer Science in the University of Edinburgh.You can find the document here.

micvog.com

Elasticsearch

Elasticsearch is a flexible and powerful open source, distributed, real time search and analytics engine. Architected from the ground up for use in distributed environments where reliability and scalability are must haves, Elasticsearch gives you the ability to move easily beyond simple full text search. Through its robust set of APIs and query DSLs, plus clients for the most popular programming languages, Elasticsearch delivers on the near limitless promises of search technology.

www.elasticsearch.org

Prediction IO

BUILD SMARTER SOFTWARE with Machine Learning

PredictionIO is an open source machine learning server for software developers to create predictive features, such as personalization, recommendation and content discovery.

prediction.io

hacks.mozilla.org

www.youtube.com

Container Cluster Manager

Kubernetes builds on top of Docker to construct a clustered container scheduling service. The goals of the project are to enable users to ask a Kubernetes cluster to run a set of containers. The system will automatically pick a worker node to run those containers on.

As container based applications and systems get larger, some tools are provided to facilitate sanity. This includes ways for containers to find and communicate with each other and ways to work with and manage sets of containers that do similar work.

When looking at the architecture of the system, we'll break it down to services that run on the worker node and services that play a "master" role.

github.com

Domino Data Labs

Domino is a platform for modern data scientists using Python, R, Matlab, and more.

Use our cloud hosted infrastructure to securely run your code on powerful hardware with a single command, without any changes to your code.

If you have your own infrastructure, our Enterprise offering provides powerful, easy to use cluster management functionality behind your firewall.

Special offer for The Machine Learning Salon's readers:

Machine Learning Salon readers can get $50 worth of compute credits when they sign up for Domino. Domino lets you run your analyses on powerful cloud hardware in one step, without any setup or changes to your code. Sign up here , or email [email protected] and tell them you are a Machine Learning Salon reader.

www.dominoup.com

Data Driven NYC Meetup Videos

Data Driven NYC is a community of tech enthusiasts who are passionate about Big Data, data technologies and data driven products and businesses, in New York and beyond. The community meets monthly at three hour events that include both company presentations and informal networking.

Data Driven NYC was founded and is organized by Matt Turck. Matt is a Managing Director at FirstMark Capital, a New York venture capital firm, where focuses on early stage technology investments in the enterprise, data, fintech, infrastructure, and connected devices sectors. In addition to Data Driven NYC, Matt founded Hardwired NYC, another community and monthly event, focused on the Internet of Things, 3D printing and wearable computing.

www.youtube.com

More to be added ...

The

Machine Learning

Salon

Download

Menu