Skip to content
Machine Learning SalonFree learning resources

Log

What is new today

A running log of new books, courses, blogs, videos and competitions.

Deep Learning Book by Yoshua Bengio, Ian Goodfellow and Aaron Courville 2014

Draft chapters available for feedback, August 2014

Please help us make this a great book! This draft is still full of typos and can be improved in many ways. Your suggestions are more than welcome. Do not hesitate to contact any of the authors directly by email or Google+ messages: Yoshua, Ian, Aaron.

www.iro.umontreal.ca

Harvard CS 171 Visualisation Course

The amount and complexity of information produced in science, engineering, business, and everyday human activity is increasing at staggering rates. The goal of this course is to expose you to visual representation methods and techniques that increase the understanding of complex data. Good visualizations not only present a visual interpretation of data, but do so by improving comprehension, communication, and decision making.

In this course you will learn how the human visual system processes and perceives images, good design practices for visualization, tools for visualization of data from a variety of fields, and programming of interactive web based visualizations using D3.

www.cs171.org

Datavu Blog

datavu.blogspot.ch

Scott Locklin's Blog

scottlocklin.wordpress.com

Barbican Digital Revolution

Digital Revolution is the most comprehensive presentation of digital creativity ever to be staged in the UK.

This immersive and interactive exhibition brings together for the first time a range of artists, filmmakers, architects, designers, musicians and game developers, all pushing the boundaries of their fields using digital media. It also looks at the dynamic developments in the areas of creative coding and DIY culture and the exciting creative possibilities offered by augmented reality, artificial intelligence, wearable technologies and 3d printing.

www.barbican.org.uk

Columbia University Applied Data Science by Ian Langmore and Daniel Krasner

The purpose of this course is to take people with strong mathematical/statistical knowledge and teach them software development fundamentals. This course will cover

• Design of small software packages

• Working in a Unix environment

• Designing software in teams

• Fundamental statistical algorithms such as linear and logistic regression

• Overfitting and how to avoid it

• Working with text data (e.g. regular expressions)

• Time series

• And more. . .

columbia applied data science.github.io

columbia applied data science.github.io

Harvard University, Data Science Course, Fall 2013

Learning from data in order to gain useful predictions and insights. This course introduces methods for five key facets of an investigation: data wrangling, cleaning, and sampling to get a suitable data set; data management to be able to access big data quickly and reliably; exploratory data analysis to generate hypotheses and intuition; prediction based on statistical methods such as regression and classification; and communication of results through visualization, stories, and interpretable summaries.

We will be using Python for all programming assignments and projects.

cm.dce.harvard.edu

MIRI

MIRI, Machine Intelligence Research Institute

The mathematics of safe machine intelligence

MIRI’s mission is to ensure that the creation of smarter than human intelligence has a positive impact. We aim to make intelligent machines behave as we intend even in the absence of immediate human supervision. Much of our current research deals with reflection, an AI’s ability to reason about its own behavior in a principled rather than ad hoc way. We focus our research on AI approaches that can be made transparent (e.g. principled decision algorithms, not genetic algorithms), so that humans can understand why the AIs behave as they do.

intelligence.org

High performance text processing in Machine Learning by Daniel Krasner

In this talk, Daniel Krasner covers rapid development of high performance scalable text processing solutions for tasks such as classification, semantic analysis, topic modeling and general machine learning. He demonstrates how Python modules, in particular the Rosetta Python library, can be used to process, clean, tokenize, extract features, and build statistical models with large volumes of text data. The Rosetta library focuses on creating small and simple modules (each with command line interfaces) that use very little memory and are parallelized with the multiprocessing package. Daniel also touches on LDA topic modeling and different implementations thereof (Vowpal Wabbit and Gensim). The talk is part presentation, and part “real life” example tutorial. This talk was recorded at the NYC Machine Learning meetup at Pivotal Labs.

www.hakkalabs.co

Data Driven NYC Meetup Videos

Data Driven NYC is a community of tech enthusiasts who are passionate about Big Data, data technologies and data driven products and businesses, in New York and beyond. The community meets monthly at three hour events that include both company presentations and informal networking.

Data Driven NYC was founded and is organized by Matt Turck. Matt is a Managing Director at FirstMark Capital, a New York venture capital firm, where focuses on early stage technology investments in the enterprise, data, fintech, infrastructure, and connected devices sectors. In addition to Data Driven NYC, Matt founded Hardwired NYC, another community and monthly event, focused on the Internet of Things, 3D printing and wearable computing.

www.youtube.com

Cisco Internet of Things Innovation Grand Challenge

The focus of the Internet of Things (IoT) Innovation Grand Challenge is to spearhead an industry wide initiative to accelerate the adoption of breakthrough technologies and products that will contribute to the growth and evolution of the Internet of Things.

This global open competition aims to recognize, promote and reward innovators, entrepreneurs and early stage startup businesses that can help us transform businesses and industries by re inventing business processes, operational efficiencies and customer service innovations.

We are seeking submissions from early stage businesses and teams that have technology based prototypes and proof of concepts (PoC) in development.

iotchallenge.cisco.spigit.com

Past, Present, and Future of Statistical Science by COPSS, 2014

nisla05.niss.org

Joseph Misiti's Blog

A curated list of awesome machine learning frameworks, libraries and software (by language). Inspired by awesome php. Other awesome lists can be found in the awesome awesomeness list.

github.com

The Machine Learning Salon's Kit has reached 100 pages!

The next milestone is set at 10,000 unique visitors (currently 3,646 unique visitors and more than 10,000 page views).

Neural Information Processing Systems Foundation (NIPS) Video resources

The Foundation: The Neural Information Processing Systems (NIPS) Foundation is a non profit corporation whose purpose is to foster the exchange of research on neural information processing systems in their biological, technological, mathematical, and theoretical aspects. Neural information processing is a field which benefits from a combined view of biological, physical, mathematical, and computational sciences.

The primary focus of the NIPS Foundation is the presentation of a continuing series of professional meetings known as the Neural Information Processing Systems Conference, held over the years at various locations in the United States, Canada and Spain.

www.youtube.com

A Few Useful Things to Know about Machine Learning, Pedro Domingos

Machine learning algorithms can figure out how to perform important tasks by generalizing from examples. This is of ten feasible and cost effective where manual programming is not. As more data becomes available, more ambitious problems can be tackled. As a result, machine learning is widely used in computer science and other fields. However, developing successful machine learning applications requires a substantial amount of “black art” that is hard to find in textbooks. This article summarizes twelve key lessons that machine learning researchers and practitioners have learned. These include pitfalls to avoid, important issues to focus on, and answers to common questions.

homes.cs.washington.edu

Gilles Louppe's Blog

Understanding Random Forest, PhD Thesis

github.com

Sebastian Raschka’s Blog

A collection of tutorials and examples for solving and understanding machine learning and pattern classification tasks

Links to useful resources

github.com

Probabilistic Programming and Bayesian Methods for Hackers by Cameron Davidson Pilon, 2014

Bayesian Methods for Hackers is designed as a introduction to Bayesian inference from a computational/understanding first, and mathematics second, point of view. Of course as an introductory book, we can only leave it at that: an introductory book. For the mathematically trained, they may cure the curiosity this text generates with other texts designed with mathematical analysis in mind. For the enthusiast with less mathematical background, or one who is not interested in the mathematics but simply the practice of Bayesian methods, this text should be sufficient and entertaining.

github.com

Carnegie Mellon University Video resources

"The videos below are intended to serve as resources for our current students, and not as online learning materials for students outside of our program.", The Machine Learning Department

www.ml.cmu.edu

Bugra Akyildiz's Blog

Great Blog (Notes) both theoretical and practical

I work as a Machine Learning/NLP Engineer at CB Insights where I apply machine learning algorithms to NLP problems. I received B.S from Bilkent University and M.Sc from New York University focusing signal processing and machine learning.

bugra.github.io

SciPy 2014

SciPy is a community dedicated to the advancement of scientific computing through open source Python software for mathematics, science, and engineering. The annual SciPy Conference allows participants from all types of organizations to showcase their latest projects, learn from skilled users and developers, and collaborate on code development.

pyvideo.org

PyLadies London Meetup resources

PyLadies is an international mentorship group with a focus on helping more women and genderqueers become active participants and leaders in the Python open source community. Our mission is to promote, educate and advance a diverse Python community through outreach, education, conferences, events, and social gatherings. PyLadies also aims to provide a friendly support network for women and genderqueers, and a bridge to the larger Python world.

github.com

Discovered on Kaggle website, a link to a very useful website:

Kaggle Competition Past Solutions

We learn more from code, and from great code. Not necessarily always the 1st ranking solution, because we also learn what makes a stellar and just a good solution. I will post solutions I came upon so we can all learn to become better!

I collected the following source code and interesting discussions from the Kaggle held competitions for learning purposes. Not all competitions are listed because I am only manually collecting them, also some competitions are not listed due to no one sharing. I will add more as time goes by. Thank you.

www.chioka.in

View on Nuit Blanche's Blog today, Lib Skylark :

The Sketching based Matrix computations for Machine Learning is a library for matrix computations suitable for general statistical data analysis and optimization applications.

Many tasks in machine learning and statistics ultimately end up being problems involving matrices: whether you're finding the key players in the bitcoin market, or inferring where tweets came from, or figuring out what's in sewage, you'll want to have a toolkit for least squares and robust regression, eigenvector analysis, non negative matrix factorization, and other matrix computations.

Sketching is a way to compress matrices that preserves key matrix properties; it can be used to speed up many matrix computations. Sketching takes a given matrix A and produces a sketch matrix B that has fewer rows and/or columns than A. For a good sketch B, if we solve a problem with input B, the solution will also be pretty good for input A. For some problems, sketches can also be used to get faster ways to find high precision solutions to the original problem. In other cases, sketches can be used to summarize the data by identifying the most important rows or columns.

A simple example of sketching is just sampling the rows (and/or columns) of the matrix, where each row (and/or column) is equally likely to be sampled. This uniform sampling is quick and easy, but doesn't always yield good sketches; however, there are sophisticated sampling methods that do yield good sketches.

xdata skylark.github.io

Starting 30 08 2014

Welcome to the Official 2014 TEXATA Big Data Analytics World Championships . This global event is a fun, innovative and challenging competition for students and professionals to develop and test their Big Data Analytics skills against their friends, colleagues and top data experts from around the world.

TEXATA 2014 is a World Championship Event independently organized and administered by the Professional Services Champions League (PSCL).

www.texata.com

Mike Bostock astonishing visualisations #machinelearning

bost.ocks.org

24/25 06 2014

www.google.com

LAST ADDITIONS

Mutual Information Text Explorer

The Mutual information Text Explorer is a tool that allows interactive exploration of text data and document covariates. See the paper or slides for information. Currently, an experimental system is available.

brenocon.com

Data Science Central

Data Science Central is the industry's online resource for big data practitioners. From Analytics to Data Integration to Visualization, Data Science Central provides a community experience that includes a robust editorial platform, social interaction, forum based technical support, the latest in technology, tools and trends and industry job opportunities.

www.datasciencecentral.com

YOU CANalytics

Welcome to UCAnalytics.com, the idea behind this website is to explore the applications of advanced Analytics and data mining in business. Analytics is an effort to explore interesting but hidden patterns in data for business growth. This idea has inspired me to name the site

• UCAnalytics: YOU CANalytics

• UCAnalytics: YOU SEE Analytics

• UCAnalytics: University for Analytics

This is sort of like finding patterns in a cluster of clouds, a fun exercise. However, we will explore some serious business applications and usage of Analytics over here. A few topics including

1. Analytical Scorecard Development

2. Customer Segmentation to gain deeper knowledge of customer behaviour

3. Data mining and Big Data Analytics

4. Business Applications of Bayesian Statistics, Nate Silver has made Bayesian cool!

5. Challenges & Pitfalls in Business Forecasting, Time Series Modelling

6. Business Growth through right Design of Experiments

7. Business Growth & Risk Estimation through Analytical simulations

Look forward to share my ideas and hear back from you.

Roopam Upadhyay

ucanalytics.com

Speaker Deck

speakerdeck.com

Code School, R Course

Learn the R programming language for data analysis and visualization. This software programming language is great for statistical computing and graphics.

www.codeschool.com

DataCamp R Course

• Introduction to R

• Data Analysis and Statistical Inference

• Introduction to Computational Finance and Financial Econometrics

• How to work with Quandl in R

www.datacamp.com

More to be added ...

The

Machine Learning

Salon

Download

Menu