30 06 2014
Deep Learning Book by Yoshua Bengio, Ian Goodfellow and Aaron Courville 2014
Draft chapters available for feedback, August 2014
Please help us make this a great book! This draft is still full of typos and can be improved in many ways. Your suggestions are more than welcome. Do not hesitate to contact any of the authors directly by email or Google+ messages: Yoshua, Ian, Aaron.
www.iro.umontreal.ca
Harvard CS 171 Visualisation Course
The amount and complexity of information produced in science, engineering, business, and everyday human activity is increasing at staggering rates. The goal of this course is to expose you to visual representation methods and techniques that increase the understanding of complex data. Good visualizations not only present a visual interpretation of data, but do so by improving comprehension, communication, and decision making.
In this course you will learn how the human visual system processes and perceives images, good design practices for visualization, tools for visualization of data from a variety of fields, and programming of interactive web based visualizations using D3.
www.cs171.org
Datavu Blog
datavu.blogspot.ch
Scott Locklin's Blog
scottlocklin.wordpress.com
Barbican Digital Revolution
Digital Revolution is the most comprehensive presentation of digital creativity ever to be staged in the UK.
This immersive and interactive exhibition brings together for the first time a range of artists, filmmakers, architects, designers, musicians and game developers, all pushing the boundaries of their fields using digital media. It also looks at the dynamic developments in the areas of creative coding and DIY culture and the exciting creative possibilities offered by augmented reality, artificial intelligence, wearable technologies and 3d printing.
www.barbican.org.uk
06 08 2014
Columbia University Applied Data Science by Ian Langmore and Daniel Krasner
The purpose of this course is to take people with strong mathematical/statistical knowledge and teach them software development fundamentals. This course will cover
• Design of small software packages
• Working in a Unix environment
• Designing software in teams
• Fundamental statistical algorithms such as linear and logistic regression
• Overfitting and how to avoid it
• Working with text data (e.g. regular expressions)
• Time series
• And more. . .
columbia applied data science.github.io
columbia applied data science.github.io
05 08 2014
Harvard University, Data Science Course, Fall 2013
Learning from data in order to gain useful predictions and insights. This course introduces methods for five key facets of an investigation: data wrangling, cleaning, and sampling to get a suitable data set; data management to be able to access big data quickly and reliably; exploratory data analysis to generate hypotheses and intuition; prediction based on statistical methods such as regression and classification; and communication of results through visualization, stories, and interpretable summaries.
We will be using Python for all programming assignments and projects.
cm.dce.harvard.edu
22 07 2014
New Blogs/Forum/Q&A in Chinese
Zhihu.com
Machine Learning
www.zhihu.com
Data Mining
www.zhihu.com
Artificial Intelligence
www.zhihu.com
Guokr.com
Machine Learning
www.guokr.com
mooc.guokr.com
Data Mining
www.guokr.com
mooc.guokr.com
Artificial Intelligence
www.guokr.com
mooc.guokr.com
20 07 2014
MIRI
MIRI, Machine Intelligence Research Institute
The mathematics of safe machine intelligence
MIRI’s mission is to ensure that the creation of smarter than human intelligence has a positive impact. We aim to make intelligent machines behave as we intend even in the absence of immediate human supervision. Much of our current research deals with reflection, an AI’s ability to reason about its own behavior in a principled rather than ad hoc way. We focus our research on AI approaches that can be made transparent (e.g. principled decision algorithms, not genetic algorithms), so that humans can understand why the AIs behave as they do.
intelligence.org
19 07 2014
High performance text processing in Machine Learning by Daniel Krasner
In this talk, Daniel Krasner covers rapid development of high performance scalable text processing solutions for tasks such as classification, semantic analysis, topic modeling and general machine learning. He demonstrates how Python modules, in particular the Rosetta Python library, can be used to process, clean, tokenize, extract features, and build statistical models with large volumes of text data. The Rosetta library focuses on creating small and simple modules (each with command line interfaces) that use very little memory and are parallelized with the multiprocessing package. Daniel also touches on LDA topic modeling and different implementations thereof (Vowpal Wabbit and Gensim). The talk is part presentation, and part “real life” example tutorial. This talk was recorded at the NYC Machine Learning meetup at Pivotal Labs.
www.hakkalabs.co
19 07 2014
Data Driven NYC Meetup Videos
Data Driven NYC is a community of tech enthusiasts who are passionate about Big Data, data technologies and data driven products and businesses, in New York and beyond. The community meets monthly at three hour events that include both company presentations and informal networking.
Data Driven NYC was founded and is organized by Matt Turck. Matt is a Managing Director at FirstMark Capital, a New York venture capital firm, where focuses on early stage technology investments in the enterprise, data, fintech, infrastructure, and connected devices sectors. In addition to Data Driven NYC, Matt founded Hardwired NYC, another community and monthly event, focused on the Internet of Things, 3D printing and wearable computing.
www.youtube.com
19 07 2014
Cisco Internet of Things Innovation Grand Challenge
The focus of the Internet of Things (IoT) Innovation Grand Challenge is to spearhead an industry wide initiative to accelerate the adoption of breakthrough technologies and products that will contribute to the growth and evolution of the Internet of Things.
This global open competition aims to recognize, promote and reward innovators, entrepreneurs and early stage startup businesses that can help us transform businesses and industries by re inventing business processes, operational efficiencies and customer service innovations.
We are seeking submissions from early stage businesses and teams that have technology based prototypes and proof of concepts (PoC) in development.
iotchallenge.cisco.spigit.com
19 07 2014
Apache Spark Summit Videos
www.youtube.com
18 07 2014
Past, Present, and Future of Statistical Science by COPSS, 2014
nisla05.niss.org
16 07 204
Joseph Misiti's Blog
A curated list of awesome machine learning frameworks, libraries and software (by language). Inspired by awesome php. Other awesome lists can be found in the awesome awesomeness list.
github.com
15 07 2014
The Machine Learning Salon's Kit has reached 100 pages!
The next milestone is set at 10,000 unique visitors (currently 3,646 unique visitors and more than 10,000 page views).
15 07 2014
Neural Information Processing Systems Foundation (NIPS) Video resources
The Foundation: The Neural Information Processing Systems (NIPS) Foundation is a non profit corporation whose purpose is to foster the exchange of research on neural information processing systems in their biological, technological, mathematical, and theoretical aspects. Neural information processing is a field which benefits from a combined view of biological, physical, mathematical, and computational sciences.
The primary focus of the NIPS Foundation is the presentation of a continuing series of professional meetings known as the Neural Information Processing Systems Conference, held over the years at various locations in the United States, Canada and Spain.
www.youtube.com
15 07 2014
A Few Useful Things to Know about Machine Learning, Pedro Domingos
Machine learning algorithms can figure out how to perform important tasks by generalizing from examples. This is of ten feasible and cost effective where manual programming is not. As more data becomes available, more ambitious problems can be tackled. As a result, machine learning is widely used in computer science and other fields. However, developing successful machine learning applications requires a substantial amount of “black art” that is hard to find in textbooks. This article summarizes twelve key lessons that machine learning researchers and practitioners have learned. These include pitfalls to avoid, important issues to focus on, and answers to common questions.
homes.cs.washington.edu
15 07 2014
Gilles Louppe's Blog
Understanding Random Forest, PhD Thesis
github.com
14 07 2014
Sebastian Raschka’s Blog
A collection of tutorials and examples for solving and understanding machine learning and pattern classification tasks
Links to useful resources
github.com
14 07 2014
Probabilistic Programming and Bayesian Methods for Hackers by Cameron Davidson Pilon, 2014
Bayesian Methods for Hackers is designed as a introduction to Bayesian inference from a computational/understanding first, and mathematics second, point of view. Of course as an introductory book, we can only leave it at that: an introductory book. For the mathematically trained, they may cure the curiosity this text generates with other texts designed with mathematical analysis in mind. For the enthusiast with less mathematical background, or one who is not interested in the mathematics but simply the practice of Bayesian methods, this text should be sufficient and entertaining.
github.com
13 07 2014
Carnegie Mellon University Video resources
"The videos below are intended to serve as resources for our current students, and not as online learning materials for students outside of our program.", The Machine Learning Department
www.ml.cmu.edu
11 07 2014
Bugra Akyildiz's Blog
Great Blog (Notes) both theoretical and practical
I work as a Machine Learning/NLP Engineer at CB Insights where I apply machine learning algorithms to NLP problems. I received B.S from Bilkent University and M.Sc from New York University focusing signal processing and machine learning.
bugra.github.io
11 07 2014
SciPy 2014
SciPy is a community dedicated to the advancement of scientific computing through open source Python software for mathematics, science, and engineering. The annual SciPy Conference allows participants from all types of organizations to showcase their latest projects, learn from skilled users and developers, and collaborate on code development.
pyvideo.org
11 07 2014
PyLadies London Meetup resources
PyLadies is an international mentorship group with a focus on helping more women and genderqueers become active participants and leaders in the Python open source community. Our mission is to promote, educate and advance a diverse Python community through outreach, education, conferences, events, and social gatherings. PyLadies also aims to provide a friendly support network for women and genderqueers, and a bridge to the larger Python world.
github.com
10 07 2014
MLSS 2014 Pittsburgh + Alex Smola's playlist
www.youtube.com
09 07 2014
Discovered on Kaggle website, a link to a very useful website:
Kaggle Competition Past Solutions
We learn more from code, and from great code. Not necessarily always the 1st ranking solution, because we also learn what makes a stellar and just a good solution. I will post solutions I came upon so we can all learn to become better!
I collected the following source code and interesting discussions from the Kaggle held competitions for learning purposes. Not all competitions are listed because I am only manually collecting them, also some competitions are not listed due to no one sharing. I will add more as time goes by. Thank you.
www.chioka.in
08 07 2014
View on Nuit Blanche's Blog today, Lib Skylark :
The Sketching based Matrix computations for Machine Learning is a library for matrix computations suitable for general statistical data analysis and optimization applications.
Many tasks in machine learning and statistics ultimately end up being problems involving matrices: whether you're finding the key players in the bitcoin market, or inferring where tweets came from, or figuring out what's in sewage, you'll want to have a toolkit for least squares and robust regression, eigenvector analysis, non negative matrix factorization, and other matrix computations.
Sketching is a way to compress matrices that preserves key matrix properties; it can be used to speed up many matrix computations. Sketching takes a given matrix A and produces a sketch matrix B that has fewer rows and/or columns than A. For a good sketch B, if we solve a problem with input B, the solution will also be pretty good for input A. For some problems, sketches can also be used to get faster ways to find high precision solutions to the original problem. In other cases, sketches can be used to summarize the data by identifying the most important rows or columns.
A simple example of sketching is just sampling the rows (and/or columns) of the matrix, where each row (and/or column) is equally likely to be sampled. This uniform sampling is quick and easy, but doesn't always yield good sketches; however, there are sophisticated sampling methods that do yield good sketches.
xdata skylark.github.io
Starting 30 08 2014
Welcome to the Official 2014 TEXATA Big Data Analytics World Championships . This global event is a fun, innovative and challenging competition for students and professionals to develop and test their Big Data Analytics skills against their friends, colleagues and top data experts from around the world.
TEXATA 2014 is a World Championship Event independently organized and administered by the Professional Services Champions League (PSCL).
www.texata.com
07 07 2014
Larry Page and Sergei Brin talking about Machine Learning, starts at 11:33 on
www.businessinsider.com
07 07 2014
Data Science related Google I/O Videos
www.youtube.com
26 06 2014
Mike Bostock astonishing visualisations #machinelearning
bost.ocks.org
24/25 06 2014
www.google.com
LAST ADDITIONS
Mutual Information Text Explorer
The Mutual information Text Explorer is a tool that allows interactive exploration of text data and document covariates. See the paper or slides for information. Currently, an experimental system is available.
brenocon.com
Data Science Central
Data Science Central is the industry's online resource for big data practitioners. From Analytics to Data Integration to Visualization, Data Science Central provides a community experience that includes a robust editorial platform, social interaction, forum based technical support, the latest in technology, tools and trends and industry job opportunities.
www.datasciencecentral.com
YOU CANalytics
Welcome to UCAnalytics.com, the idea behind this website is to explore the applications of advanced Analytics and data mining in business. Analytics is an effort to explore interesting but hidden patterns in data for business growth. This idea has inspired me to name the site
• UCAnalytics: YOU CANalytics
• UCAnalytics: YOU SEE Analytics
• UCAnalytics: University for Analytics
This is sort of like finding patterns in a cluster of clouds, a fun exercise. However, we will explore some serious business applications and usage of Analytics over here. A few topics including
1. Analytical Scorecard Development
2. Customer Segmentation to gain deeper knowledge of customer behaviour
3. Data mining and Big Data Analytics
4. Business Applications of Bayesian Statistics, Nate Silver has made Bayesian cool!
5. Challenges & Pitfalls in Business Forecasting, Time Series Modelling
6. Business Growth through right Design of Experiments
7. Business Growth & Risk Estimation through Analytical simulations
Look forward to share my ideas and hear back from you.
Roopam Upadhyay
ucanalytics.com
Speaker Deck
speakerdeck.com
Code School, R Course
Learn the R programming language for data analysis and visualization. This software programming language is great for statistical computing and graphics.
www.codeschool.com
DataCamp R Course
• Introduction to R
• Data Analysis and Statistical Inference
• Introduction to Computational Finance and Financial Econometrics
• How to work with Quandl in R
www.datacamp.com
More to be added ...
The
Machine Learning
Salon
Download
Menu