
Clustering with DBSCAN
DBSCAN is a density-based clustering algorithm for doing unsupervised learning. It's pretty nifty: with just two parameters, you can specify "dense" regions in your data, and grow those regions out o...
20 Nov 201716min

The Kaggle Survey on Data Science
Want to know what's going on in data science these days? There's no better way than to analyze a survey with over 16,000 responses that recently released by Kaggle. Kaggle asked practicing and aspir...
13 Nov 201725min

Machine Learning: The High Interest Credit Card of Technical Debt
This week, we've got a fun paper by our friends at Google about the hidden costs of maintaining machine learning workflows. If you've worked in software before, you're probably familiar with the idea...
6 Nov 201722min

Improving Upon a First-Draft Data Science Analysis
There are a lot of good resources out there for getting started with data science and machine learning, where you can walk through starting with a dataset and ending up with a model and set of predict...
30 Okt 201715min

Survey Raking
It's quite common for survey respondents not to be representative of the larger population from which they are drawn. But if you're a researcher, you need to study the larger population using data fr...
23 Okt 201717min

Happy Hacktoberfest
It's the middle of October, so you've already made two pull requests to open source repos, right? If you have no idea what we're talking about, spend the next 20 minutes or so with us talking about th...
16 Okt 201715min

Re - Release: Kalman Runners
In honor of the Chicago marathon this weekend (and due in large part to Katie recovering from running in it...) we have a re-release of an episode about Kalman filters, which is part algorithm part el...
9 Okt 201717min

Neural Net Dropout
Neural networks are complex models with many parameters and can be prone to overfitting. There's a surprisingly simple way to guard against this: randomly destroy connections between hidden units, al...
2 Okt 201718min




















