31 10 2014
BerkeleyISchool Videos
The School of Information is both UC Berkeley's newest and its smallest school. Located in the center of campus, the I School is a graduate education and research community that crosses the boundaries of law, economics, sociology, business, library science, engineering, design, publishing, linguistics, computer science, and information science.
In UC Berkeley's historic South Hall, roughly 100 graduate students and 13 faculty members form a small, multidisciplinary collective of scholars and practitioners.
The I School offers a professional master's degree and an academic doctoral degree. Our master's program trains students for careers as information professionals and emphasizes small classes and project based learning. Our Ph.D. program equips scholars to develop solutions and shape policies that influence how people seek, use, and share information.
Playlists
12 VIDEOS
DataEDGE 2014
3 months ago
6 VIDEOS
Privacy & Big Data Workshop: Values and Governance
7 months ago
6 VIDEOS
Admissions FAQ
1 year ago
10 VIDEOS
DataEDGE Conference 2013
1 year ago
15 VIDEOS
Analyzing Big Data with Twitter (Info 290, Fall 2012)
2 years ago
10 VIDEOS
Social Data Revolution (Info 290A, Fall 2012)
2 years ago
11 VIDEOS
DataEDGE Conference 2012
2 years ago
Links before 24 oct 2014
Coursera: Mining Massive Datasets
We introduce the student to modern distributed file systems and MapReduce, including what distinguishes good MapReduce algorithms from good algorithms in general. The rest of the course is devoted to algorithms for extracting models and information from large datasets. Students will learn how Google's PageRank algorithm models importance of Web pages and some of the many extensions that have been used for a variety of purposes. We'll cover locality sensitive hashing, a bit of magic that allows you to find similar items in a set of items so large you cannot possibly compare each pair. When data is stored as a very large, sparse matrix, dimensionality reduction is often a good way to model the data, but standard approaches do not scale well; we'll talk about efficient approaches. Many other large scale algorithms are covered as well, as outlined in the course syllabus.
Course Syllabus
Week 1:
MapReduce
Link Analysis PageRank
Week 2:
Locality Sensitive Hashing Basics + Applications
Distance Measures
Nearest Neighbors
Frequent Itemsets
Week 3:
Data Stream Mining
Analysis of Large Graphs
Week 4:
Recommender Systems
Dimensionality Reduction
Week 5:
Clustering
Computational Advertising
Week 6:
support vector Machines
Decision Trees
MapReduce Algorithms
Week 7:
More About Link Analysis Topic specific PageRank, Link Spam.
More About Locality Sensitive Hashing
The Caltech JPL Summer School on Big Data Analytics
The anticipated schedule of lectures (subject to changes):
Each bullet bellow corresponds to a set of materials that includes approximately 2 hours of video lectures, various links and supplementary materials, plus some online, hands on exercises.
1. Introduction to the school. Software architectures. Introduction to Machine Learning.
2. Best programming practices. Information retrieval.
3. Introduction to R. Markov Chain Monte Carlo.
4. Statistical resampling and inference.
5. Databases.
6. Data visualization.
7. Clustering and classification.
8. Decision trees and random forests.
9. Dimensionality reduction. Closing remarks.
Apache Spark Machine Learning Library
MLlib is a Spark implementation of some common machine learning (ML) functionality, as well associated tests and data generators. MLlib currently supports four common types of machine learning problem settings, namely, binary classification, regression, clustering and collaborative filtering, as well as an underlying gradient descent optimization primitive.
Apache Spark Summit Videos, 2014
Apache Spark Summit exercises, 2013
Welcome to the Spark Summit hands on exercises. These exercises are adapted from similar exercises that were prepared for and run at AMP Camp Big Data Bootcamps. They were written by volunteer graduate students and postdocs in the UC Berkeley AMPLab. Many of those same graduate students are also volunteers here on the Spark Summit Training day team as well. The exercises we cover today will have you working directly with the Spark specific components of the AMPLab’s open source software stack, called the Berkeley Data Analytics Stack (BDAS).
Apache Spark Summit Training, 2014
Course Prerequisites:
Laptop with WiFi capabilities
Java 6 or 7
TRACK A: Introduction to Apache Spark Workshop
INTRO EXERCISES
The Introduction to Apache Spark workshop is for users to learn the core Spark APIs. This session features hands on technical exercises to get developers up to speed in using Spark for data exploration, analysis, and building big data applications.
The integrated lecture and lab format covers the following topics:
Overview of Big Data and Spark
Installing Spark Locally
Using Spark’s Core APIs in Scala, Java, & Python
Building Spark Applications
Deploying on a Big Data Cluster
Building Applications for Multiple Platforms
TRACK B:Advanced Apache Spark Workshop
ADVANCED EXERCISES
The Advanced Apache Spark Workshop will cover advanced topics on architecture, tuning, and each of Spark’s high level libraries (including the latest features). Attendees will have the opportunity after the lunch break to work through labs on each of the libraries.
Some familiarity with Spark or MapReduce is expected, as this workshop will not cover basic Spark programming.
Topics covered include:
Advanced Spark Internals and Tuning, Reynold Xin, SLIDES
Spark SQL, Michael Armburst, SLIDES
Spark Streaming, Tathagata Das, SLIDES
MLlib, Ameet Talwalkar, SLIDES
GraphX, Ankur Dave, SLIDES
Apache Mahout ML library
The Apache Mahout™ project's goal is to build a scalable machine learning library.
Currently Mahout supports mainly three use cases: Recommendation mining takes users' behavior and from that tries to find items users might like. Clustering takes e.g. text documents and groups them into groups of topically related documents. Classification learns from exisiting categorized documents what documents of a specific category look like and is able to assign unlabelled documents to the (hopefully) correct category.
Apache Mahout on Javaworld
Enjoy machine learning with Mahout on Hadoop, 2014
Mahout brings the power of scalable processing to Hadoop's huge data sets
Know this right now about Hadoop, 2014
From core elements like HDFS and YARN to ancillary tools like Zookeeper, Flume, and Sqoop, here's your cheat sheet and cartography of the ever expanding Hadoop ecosystem.
MapReduce programming with Apache Hadoop, 2008
Process massive data sets in parallel on large clusters
Deeplearning4j
Deeplearning4j is the first commercial grade deep learning library written in Java. It is meant to be used in business environments, rather than as a research tool for extensive data exploration. Deeplearning4j is most helpful in solving distinct problems, like identifying faces, voices, spam or e commerce fraud.
Deeplearning4j aims to be cutting edge plug and play, more convention than configuration. By following its conventions, you get an infinitely scalable deep learning architecture. The framework has a domain specific language (DSL) for neural networks, to turn their multiple knobs.
Deeplearning4j includes a distributed deep learning framework and a normal deep learning framework; i.e. it runs on a single thread as well. Training takes place in the cluster, which means it can process massive amounts of data. Nets are trained in parallel via iterative reduce.
The distributed framework is made for data input and neural net training at scale, and its output should be highly accurate predictive models.
By following the links at the bottom of each page, you will learn to set up, and train with sample data, several types of deep learning networks. These include single and multithread networks, Restricted Boltzmann machines, deep belief networks and Stacked Denoising Autoencoders.
For a quick introduction to neural nets, please see our overview.
Udacity opencourseware "Intro to Hadoop and MapReduce"
Course Summary
The Apache™ Hadoop® project develops open source software for reliable, scalable, distributed computing. Learn the fundamental principles behind it, and how you can use its power to make sense of your Big Data.
Why Take This Course?
• How Hadoop fits into the world (recognize the problems it solves)
• Understand the concepts of HDFS and MapReduce (find out how it solves the problems)
• Write MapReduce programs (see how we solve the problems)
• Practice solving problems on your own
Storm Apache
Apache Storm is a free and open source distributed realtime computation system. Storm makes it easy to reliably process unbounded streams of data, doing for realtime processing what Hadoop did for batch processing. Storm is simple, can be used with any programming language, and is a lot of fun to use!
Michael Viogiatzis Blog
How to spot first stories on Twitter using Storm
As a first blog post, I decided to describe a way to detect first stories (a.k.a new events) on Twitter as they happen. This work is part of the Thesis I wrote last year for my MSc in Computer Science in the University of Edinburgh.You can find the document here.
Elasticsearch
Elasticsearch is a flexible and powerful open source, distributed, real time search and analytics engine. Architected from the ground up for use in distributed environments where reliability and scalability are must haves, Elasticsearch gives you the ability to move easily beyond simple full text search. Through its robust set of APIs and query DSLs, plus clients for the most popular programming languages, Elasticsearch delivers on the near limitless promises of search technology.
Prediction IO
BUILD SMARTER SOFTWARE with Machine Learning
PredictionIO is an open source machine learning server for software developers to create predictive features, such as personalization, recommendation and content discovery.
Container Cluster Manager
Kubernetes builds on top of Docker to construct a clustered container scheduling service. The goals of the project are to enable users to ask a Kubernetes cluster to run a set of containers. The system will automatically pick a worker node to run those containers on.
As container based applications and systems get larger, some tools are provided to facilitate sanity. This includes ways for containers to find and communicate with each other and ways to work with and manage sets of containers that do similar work.
When looking at the architecture of the system, we'll break it down to services that run on the worker node and services that play a "master" role.
Domino Data Labs
Domino is a platform for modern data scientists using Python, R, Matlab, and more.
Use our cloud hosted infrastructure to securely run your code on powerful hardware with a single command, without any changes to your code.
If you have your own infrastructure, our Enterprise offering provides powerful, easy to use cluster management functionality behind your firewall.
Special offer for The Machine Learning Salon's readers:
Machine Learning Salon readers can get $50 worth of compute credits when they sign up for Domino. Domino lets you run your analyses on powerful cloud hardware in one step, without any setup or changes to your code. Sign up here , or email [email protected] and tell them you are a Machine Learning Salon reader.
Data Driven NYC Meetup Videos
Data Driven NYC is a community of tech enthusiasts who are passionate about Big Data, data technologies and data driven products and businesses, in New York and beyond. The community meets monthly at three hour events that include both company presentations and informal networking.
Data Driven NYC was founded and is organized by Matt Turck. Matt is a Managing Director at FirstMark Capital, a New York venture capital firm, where focuses on early stage technology investments in the enterprise, data, fintech, infrastructure, and connected devices sectors. In addition to Data Driven NYC, Matt founded Hardwired NYC, another community and monthly event, focused on the Internet of Things, 3D printing and wearable computing.
More to be added ...
The
Machine Learning
Salon
Download
Menu