Skip to content
Machine Learning SalonFree learning resources

Data

Open datasets

Public data collections for training models and running experiments.

CrowFlower Open Data Library

CrowdFlower encourages developers and researchers to use its open data to explore new ways of what crowdsourcing can achieve. This webpage is a repository of data sets collected or enhanced by CrowdFlower's workforce and made available for everyone to use.

www.crowdflower.com

The Million Song Dataset

The Million Song Dataset is a freely available collection of audio features and metadata for a million contemporary popular music tracks.

Its purposes are:

To encourage research on algorithms that scale to commercial sizes

To provide a reference dataset for evaluating research

As a shortcut alternative to creating a large dataset with APIs (e.g. The Echo Nest's)

To help new researchers get started in the MIR field

labrosa.ee.columbia.edu

Awesome Public Datasets by Xiaming Chen, Shanghai, China

This list of public data sources are collected and tidyed from blogs, answers, and user reponses. Most of the data sets listed below are free, however, some are not.

github.com

I am now a Ph.D. candidate with Prof. Yaohui Jin at Shanghai Jiao Tong Univ.. I received my B.S. (2010) of Optical Information and Science Technology at Xidian University, Xi'an, China.

My research interests come from the measurement and analysis of network traffic, especially the renewed models and characteristics of networks traffic, with the data mining techniques and high performance processing platforms like Network Processors and distributed processing systems like Hadoop/MapReduce or Spark.

If you are interested in my articles, researches, or projects, you can reach me via email or other partially instant messages like github.

Enjoy! : )

hsiamin.com

Google Public Data Explorer

The Google Public Data Explorer makes large, public interest datasets easy to explore, visualize and communicate. As the charts and maps animate over time, the changes in the world become easier to understand. You don't have to be a data expert to navigate between different views, make your own comparisons, and share your findings.

Students, journalists, policy makers and everyone else can play with the tool to create visualizations of public data, link to them, or embed them in their own webpages. Embedded charts and links can update automatically so you’re always sharing the latest available data.

The Public Data Explorer launched in March, 2010. See this blog post, which originally announced the product, for more background and historical perspective.

www.google.com

Gaussian Processes List of Datasets

Welcome to the web site for theory and applications of Gaussian Processes

Gaussian Process is powerful non parametric machine learning technique for constructing comprehensive probabilistic models of real world problems. They can be applied to geostatistics, supervised, unsupervised, reinforcement learning, principal component analysis, system identification and control, rendering music performance, optimization and many other tasks.

People

Geology & Modelling Research Group at Rio Tinto Centre for Mine Automation, ACFR, University of Sydney

gaussianprocess.com

The Text REtrieval Conference (TREC) Datasets

The Text REtrieval Conference (TREC), co sponsored by the National Institute of Standards and Technology (NIST) and U.S. Department of Defense, was started in 1992 as part of the TIPSTER Text program. Its purpose was to support research within the information retrieval community by providing the infrastructure necessary for large scale evaluation of text retrieval methodologies. In particular, the TREC workshop series has the following goals:

to encourage research in information retrieval based on large test collections;

to increase communication among industry, academia, and government by creating an open forum for the exchange of research ideas;

to speed the transfer of technology from research labs into commercial products by demonstrating substantial improvements in retrieval methodologies on real world problems; and

to increase the availability of appropriate evaluation techniques for use by industry and academia, including development of new evaluation techniques more applicable to current systems.

TREC is overseen by a program committee consisting of representatives from government, industry, and academia. For each TREC, NIST provides a test set of documents and questions. Participants run their own retrieval systems on the data, and return to NIST a list of the retrieved top ranked documents. NIST pools the individual results, judges the retrieved documents for correctness, and evaluates the results. The TREC cycle ends with a workshop that is a forum for participants to share their experiences.

trec.nist.gov

trec.nist.gov

The New York Times Linked Open Data (Beta)

For the last 150 years, The New York Times has maintained one of the most authoritative news vocabularies ever developed. In 2009, we began to publish this vocabulary as linked open data.

The Data

As of 13 January 2010, The New York Times has published approximately ,10,000 subject headings as linked open data under a CC BY license. We provide both RDF documents and a human friendly HTML versions. The table below gives a breakdown of the various tag types and mapping strategies on data.nytimes.com.

Type Manually Mapped Tags Automatically Mapped Tags Total

People 4,978 0 4,978

Organizations 1,489 1,592 3,081

Locations 1,910 0 1,910

Descriptors 498 0 498

Total 10,467

data.nytimes.com

Google Books Ngrams

Public Data Sets>Google Books Ngrams

A data set containing Google Books n gram corpuses. This data set is freely available on Amazon S3 in a Hadoop friendly file format and is licensed under a Creative Commons Attribution 3.0 Unported License. The original dataset is available from http://books.google.com/ngrams/.

aws.amazon.com

books.google.com

IMAGENET

ImageNet is an image database organized according to the WordNet hierarchy (currently only the nouns), in which each node of the hierarchy is depicted by hundreds and thousands of images. Currently we have an average of over five hundred images per node. We hope ImageNet will become a useful resource for researchers, educators, students and all of you who share our passion for pictures.

Who uses ImageNet?

We envision ImageNet as a useful resource to researchers in the academic world, as well as educators around the world.

Does ImageNet own the images? Can I download the images?

No, ImageNet does not own the copyright of the images. ImageNet only provides thumbnails and URLs of images, in a way similar to what image search engines do. In other words, ImageNet compiles an accurate list of web images for each synset of WordNet. For researchers and educators who wish to use the images for non commercial research and/or educational purposes, we can provide access through our site under certain conditions and terms. For details click here

www.image net.org

HDX Humanitarian Data Exchange

What is HDX?

The goal of the Humanitarian Data Exchange (HDX) is to make humanitarian data easy to find and use for analysis. We are working on three elements that will eventually combine into an integrated data platform.

Repository

The HDX repository, where data providers can upload their raw data spreadsheets for others to find and use.

Analytics

HDX analytics, a database of high value data that can be compared across countries and crises, with tools for analysis and visualisation.

Standards

Standards to help share humanitarian data through the use of a consensus Humanitarian Exchange Language.

data.hdx.rwlabs.org

L’Open Data français cartographié

Voici trois cartographies de l’écosphère de l‘Open Data français. Sur fond noir, les trois posters (téléchargeable au format « A0″) livrent un aperçu général sur l’open data français actuel. Les trois cartographies sont basées sur les données fournies par Data Publica, notamment deux études réalisées récemment par Guillaume Lebourgeois, Pierrick Boitel et Perrine Letellier (ayant accueilli les deux derniers dans mon enseignement à l’UTC au semestre dernier). L’objectif de ces cartes est d’entamer une « radiographie » assez complète du domaine, renouvelable dans le temps (peut être tous les six mois) et directement associée aux données présentes chez Data Publica. En somme, une sorte d’observatoire de l’open data français dans lequel je me lance à travers les productions de l’Atelier de Cartographie.

ateliercartographie.wordpress.com

DATA SOURCES FOR COOL DATA SCIENCE PROJECTS: PART 1 / 2, GUEST POST

At The Data Incubator, we run a free six week data science fellowship to help our Fellows land industry jobs. Our hiring partners love considering Fellows who don’t mind getting their hands dirty with data. That’s why our Fellows work on cool capstone projects that showcase those skills. One of the biggest obstacles to successful projects has been getting access to interesting data. Here are a few cool public data sources you can use for your next project:

Economic Data:

Publically Traded Market Data: Quandl is an amazing source of finance data. Google Finance and Yahoo Finance are additional good sources of data. Corporate filings with the SEC are available on Edgar.

Housing Price Data: You can use the Trulia API or the Zillow API.

Lending data: You can find student loan defaults by university and the complete collection of peer to peer loans from Lending Club and Prosper, the two largest platforms in the space.

Home mortgage data: There is data made available by the Home Mortgage Disclosure Act and there’s a lot of data from the Federal Housing Finance Agency available here.

Content Data:

Review Content: You can get reviews of restaurant and physical venues from Foursquare and Yelp (see geodata). Amazon has a large repository of Product Reviews. Beer reviews from Beer Advocate can be found here. Rotten Tomatoes Movie Reviews are available from Kaggle.

Web Content: Looking for web content? Wikipedia provides dumps of their articles. Common Crawl has a large corpus of the internet available. ArXiv maintains all their data available via Bulk Download from AWS S3. Want to know which URLs are malicious? There’s a dataset for that. Music data is available from the Million Songs Database. You can analyze the Q&A patterns on sites like Stack Exchange (including Stack Overflow).

Media Data: There’s open annotated articles form the New York Times, Reuters Dataset, and GDELT project (a consolidation of many different news sources). Google Books has published NGrams for books going back to past 1800.

Communications Data: There’s access to public messages of the Apache Software Foundation and communications amongst former execs Enron

Government Data:

Municipal Data: Crime Data is available for City of Chicago, and Washington DC. Restaurant Inspection Data is available for Chicago and New York City.

Transportation Data: NYC Taxi Trips in 2013 are available courtesy of the Freedom of Information Act. There’s bikesharing data from NYC, Washington DC, and SF. There’s also Flight Delay Data from the FAA

Census Data: Japanese Census Data. US Census data from 2010, 2000, 1990. From census data, the government has also derived time use data. EU Census Data. Checkout popular male / female baby names going back to the 19th Century from the Social Security Administration.

World Bank: they have a lot of data available on their website.

Election Data: Political contribution data for the last few US elections can be downloaded from the FEC here and here. Polling data is available from Real Clear Politics.

Data With a Cause:

Environmental Data: Data on household energy usage is available as well as NASA Climate Data.

Medical and biological Data: You can get anything from anonymous medical records, to remote sensor reading for individuals, to data of the Genomes of 1000 individuals.

Miscellaneous:

Geo Data: Try looking at these Yelp Datasets for venues near major universities and one for major cities in the Southwest. The Foursquare API is another good source. Open Street Map has open data on venues as well.

Twitter Data: you can get access to Twitter Data used for sentiment analysis, network Twitter Data, social Twitter data, on top of their API.

Games Data: Datasets for games, including a large dataset of Poker hands, dataset of online Domion Games, and datasets of Chess Games are available.

Web Usage Data: Web usage data is a common dataset that companies look at to understand engagement. Available datasets include Anonymous usage data for MSNBC, Amazon purchase history (also anonymized), and Wikipedia traffic.

Metasources: these are great sources for other web pages.

Stanford Network Data: http://snap.stanford.edu/index.html

Every year, the ACM holds a competition for machine learning called the KDD Cup. Their data is available online.

UCI maintains archives of data for machine learning.

US Census Data

Amazon is hosting Public Datasets on s3

Kaggle hosts machine learning challenges and many of their datasets are publicly available

The cities of Chicago, New York, Washington DC, and SF maintain public data warehouses.

Yahoo maintains a lot of data on its web properties which can be obtained by writing them.

BigML is a blog that maintains a list of public datasets for the machine learning community.

Finally, if there’s a website with data you are interested in, crawl for it!

101.datascience.community

101.datascience.community

California Department of Water Resources

DWR has many programs and data tools to collect and disseminate information on water resources.

All Water Data Topics… http://www.water.ca.gov/nav/nav.cfm?loc=t&id=106

CALIFORNIA DATA EXCHANGE CENTER (CDEC)

With the cooperation of over 140 other agencies, the CDEC provides real time, forecast, and historical hydrologic data. This data includes water discharge in rivers, water storage in reservoirs, precipitation accumulation, and water content in snow pack, primarily focused in flood management. However, the data is also helpful for determining general water availability and natural supply trends.

More about CDEC http://cdec.water.ca.gov/

CALIFORNIA IRRIGATION MANAGEMENT INFORMATION SYSTEM (CIMIS)

CIMIS is a network of over 120 automated weather stations in California. CIMIS was developed in 1982 by DWR and the University of California, Davis to assist California's irrigators to manage their water resources efficiently.

More about CIMIS http://wwwcimis.water.ca.gov/cimis/welcome.jsp

WATER DATA LIBRARY

The library provides geographic based data on water conditions.

More about the Water Data Library http://www.water.ca.gov/waterdatalibrary/

INTERAGENCY ECOLOGICAL PROGRAM

The Interagency Ecological Program (IEP) provides ecological information and scientific leadership for use in management of the San Francisco Estuary.

More about IEP http://www.water.ca.gov/iep/

INTEGRATED WATER RESOURCES INFORMATION SYSTEM (IWRIS)

IWRIS is a one stop shop for state wide water resources information. It integrates multidisciplinary data to support Integrated Regional Water Management.

More about IWRIS http://www.water.ca.gov/iwris/

www.water.ca.gov

LONDON DATASTORE, 591 datasets

Welcome to the new look DataStore

Over the last few months we have been busy updating London Datastore to deliver a host of practical new features, improved (geography based) searches, dataset previews and APIs, all of which will make for a much sleeker experience. The technical improvements are there to support our broader aim of kick starting collaboration so that the value of data in our city reaches its full potential.

Have a look around, read the introductory blog and Let us know what you think.

data.london.gov.uk

Links before 24 oct 2014

World Data Bank

Explore. Create. Share: Development Data

DataBank is an analysis and visualisation tool that contains collections of time series data on a variety of topics. You can create your own queries; generate tables, charts, and maps; and easily save, embed, and share them.

The World Bank Group has set two goals for the world to achieve by 2030:

• End extreme poverty by decreasing the percentage of people living on less than $1.25 a day to no more than 3%

• Promote shared prosperity by fostering the income growth of the bottom 40% for every country

The World Bank is a vital source of financial and technical assistance to developing countries around the world. We are not a bank in the ordinary sense but a unique partnership to reduce poverty and support development. The World Bank Group comprises five institutions managed by their member countries.

Established in 1944, the World Bank Group is headquartered in Washington, D.C. We have more than 10,000 employees in more than 120 offices worldwide.

databank.worldbank.org

US Dataset

The home of the U.S. Government’s open data

Here you will find data, tools, and resources to conduct research, develop web and mobile applications, design data visualizations, and more.

www.data.gov

US City Open Data Census

us city.census.okfn.org

UK Dataset

Opening up government

data.gov.uk

Machine Learning repository

The UCI Machine Learning Repository is a collection of databases, domain theories, and data generators that are used by the machine learning community for the empirical analysis of machine learning algorithms. The archive was created as an ftp archive in 1987 by David Aha and fellow graduate students at UC Irvine. Since that time, it has been widely used by students, educators, and researchers all over the world as a primary source of machine learning data sets. As an indication of the impact of the archive, it has been cited over 1000 times, making it one of the top 100 most cited "papers" in all of computer science. The current version of the web site was designed in 2007 by Arthur Asuncion and David Newman, and this project is in collaboration with Rexa.info at the University of Massachusetts Amherst. Funding support from the National Science Foundation is gratefully acknowledged.

archive.ics.uci.edu

Stanford Large Network Dataset Collection

Social networks : online social networks, edges represent interactions between people

Networks with ground truth communities : ground truth network communities in social and information networks

Communication networks : email communication networks with edges representing communication

Citation networks : nodes represent papers, edges represent citations

Collaboration networks : nodes represent scientists, edges represent collaborations (co authoring a paper)

Web graphs : nodes represent webpages and edges are hyperlinks

Amazon networks : nodes represent products and edges link commonly co purchased products

Internet networks : nodes represent computers and edges communication

Road networks : nodes represent intersections and edges roads connecting the intersections

Autonomous systems : graphs of the internet

Signed networks : networks with positive and negative edges (friend/foe, trust/distrust)

Location based online social networks : Social networks with geographic check ins

Wikipedia networks and metadata : Talk, editing and voting data from Wikipedia

Twitter and Memetracker : Memetracker phrases, links and 467 million Tweets

Online communities : Data from online communities such as Reddit and Flickr

Online reviews : Data from online review systems such as BeerAdvocate and Amazon

Information cascades : ...

SNAP networks are also availalbe from UF Sparse Matrix collection. Visualizations of SNAP networks by Tim Davis.

snap.stanford.edu

Deep Learning datasets

Deep Learning is a new area of Machine Learning research, which has been introduced with the objective of moving Machine Learning closer to one of its original goals: Artificial Intelligence.

This website is intended to host a variety of resources and pointers to information about Deep Learning. In these pages you will find

• a reading list,

• links to software,

• datasets,

• a list of deep learning research groups and labs,

• a list of announcements for deep learning related jobs (job listings),

• as well as tutorials and cool demos.

deeplearning.net

Open Government Data (OGD) Platform India

data.gov.in

Yahoo Datasets

We have various types of data available to share. They are categorized into Ratings, Language, Graph, Advertising and Market Data, Computing Systems and an appendix of other relevant data and resources available via the Yahoo! Developer Network.

webscope.sandbox.yahoo.com

Windows Azure Marketplace

One Stop Shop for Premium Data and Applications

Hundreds of Apps, Thousands of Subscriptions, Trillions of Data Points

datamarket.azure.com

Amazon Public Data Sets

Public Data Sets on AWS provides a centralized repository of public data sets that can be seamlessly integrated into AWS cloud based applications. AWS is hosting the public data sets at no charge for the community, and like all AWS services, users pay only for the compute and storage they use for their own applications. Learn more about Public Data Sets on AWS and visit the Public Data Sets forum.

aws.amazon.com

Wikipedia: Database Download

Wikipedia offers free copies of all available content to interested users. These databases can be used for mirroring, personal use, informal backups, offline use or database queries (such as for Wikipedia:Maintenance). All text content is multi licensed under the Creative Commons Attribution ShareAlike 3.0 License (CC BY SA) and the GNU Free Documentation License (GFDL). Images and other files are available under different terms, as detailed on their description pages. For our advice about complying with these licenses, see Wikipedia:Copyrights.

en.wikipedia.org

Gutenberg project (Free books available in different format, useful for NLP)

Project Gutenberg offers 45,541 free ebooks to download. (source the 5th June 2014)

www.gutenberg.org

Freebase

Use Freebase data

Freebase data is free to use under an open license. You can:

Query Freebase using our Search, Topic, or MQL APIs

Download our weekly data dumps

www.freebase.com

Datamob Data

datamob.org

Reddit Datasets

www.reddit.com

100+ Interesting Data Sets for Statistics

Summary: Looking for interesting data sets? Here's a list of more than 100 of the best stuff, from dolphin relationships to political campaign donations to death row prisoners.

rs.io

Data portal of the City of Chicago

data.cityofchicago.org

Remark: you need to copy the following link in your browser, temporary problem

Gold mine where we can find data set such as names, salaries, positions of all persons working for Chicago City!

data.cityofchicago.org

Data portal of the City of Seattle

data.seattle.gov

Data portal of the City of LA

data.lacity.org

Remark: you need to copy the following link in your browser, temporary problem

Data portal of the City of Dallas

www.dallasopendata.com

Data portal of the City of Austin

data.austintexas.gov

How to produce and use datasets: lessons learned, mlwave

mlwave.com

MITx and HarvardX release MOOC datasets and visualization tools

newsoffice.mit.edu

Transport For London Open Data, UK

www.tfl.gov.uk

Finding the perfect house using open data, Justin Palmer’s Blog

dealloc.me

Synapse

A private or public workspace that allows you to aggregate, describe, and share your research.

A tool to improve reproducibility of data intensive science, recording progress as you work with tools such as R and Python.

A set of living research projects enabling contribution to large scale collaborative solutions to scientific problems.

www.synapse.org

Montreal, Portail Donnees Ouvertes (French&English), Canada

donnees.ville.montreal.qc.ca

Insee, France

www.insee.fr

RATP Open Data, French Tube in Paris, France

data.ratp.fr

The

Machine Learning

Salon

Download

Menu