Category: Data Science (Page 2 of 3)

Pandas: Filtering and segmenting

December 1, 2020 / Brett Romero

This article is part of a series of practical guides for using the Python data processing library pandas. To see view all the available parts, click here.

One of the most common ways you will interact with a pandas DataFrame is by selecting different combinations of columns and rows. This can be done using the numerical positions of columns and rows in the DataFrame, column names and row indices, or by filtering the rows by applying some criteria to the data in the DataFrame. All of these options (and combinations of them) are available, so let’s dig in!

Pandas: Basic data interrogation

November 24, 2020 / Brett Romero

This article is part of a series of practical guides for using the Python data processing library pandas. To see view all the available parts, click here.

Once we have our data in a pandas DataFrame, the basic table structure in pandas, the next step is how do we assess what we have? If you are coming from Excel or R Studio, you are probably used to being able to look at the data any time you want. In python/pandas, we don’t have a spreadsheet to work with, and we don’t even have an equivalent of R Studio (although Jupyter notebooks are a similar concept), but we do have several tools available that can help you get a handle on what your data looks like.

Pandas: Reading in JSON data

November 23, 2020 / Brett Romero

This article is part of a series of practical guides for using the Python data processing library pandas. To see view all the available parts, click here.

When we are working with data in software development or when the data comes from APIs, it is often not provided in a tabular form. Instead it is provided in some combination of key-value stores and arrays broadly denoted as JavaScript Object Notation (JSON). So how do we read this type of non-tabular data into a tabular format like a pandas DataFrame?

Pandas: Reading in tabular data

November 13, 2020 / Brett Romero

This article is part of a series of practical guides for using the Python data processing library pandas. To see view all the available parts, click here.

To get started with pandas, the first thing you are going to need to understand is how to get data into pandas. For this guide we are going to focus on reading in tabular data (i.e. data stored in a table with rows and columns). If you don’t have some data available but want to try some things out, a great place to get some data to play with is the UCI Machine Learning Repository.

Why the ‘boring’ part of Data Science is actually the most interesting

May 30, 2017 / Brett Romero / 1 Comment

For the last 5 years, data science has been one of the world’s hottest professions, but it is also one of the most poorly defined. This can be seen on any career website, where advertisements for ‘Data Scientist’ positions describe everything from what used to be a simple data analyst role, to technical, PhD-only, research positions working on artificial intelligence or autonomous cars.

Forget SQL or NoSQL – 5 scenarios where you may not need a database at all

April 18, 2017 / Brett Romero / 0 Comments

A while back, I attended a hackathon in Belgrade as a mentor. This hackathon was the first ‘open data’ hackathon in Serbia and focused on making applications using data that had recently been released by various ministries, government agencies, and independent bodies in Serbia. As we walked around talking to the various teams, one of the things I noticed at the time, was that almost all teams were using databases to manage their data . In most cases, the database being used was something very lightweight like SQLite3, but in some cases more serious databases (MySQL, PostgreSQL, MongoDB) were also being used.

Data Science: A Kaggle Walkthrough – Creating a Model

May 8, 2016 / Brett Romero / 0 Comments

This article is Part VI in a series looking at data science and machine learning by walking through a Kaggle competition. If you have not done so already, you are strongly encouraged to go back and read the earlier parts – (Part I, Part II, Part III, Part IV and Part V).

Continuing on the walkthrough, in this part we build the model that will predict the first booking destination country for each user based on the dataset created in the earlier parts.

Data Science: A Kaggle Walkthrough – Adding New Data

April 12, 2016 / Brett Romero / 0 Comments

This article is Part V in a series looking at data science and machine learning by walking through a Kaggle competition. If you have not done so already, you are strongly encouraged to go back and read the earlier parts – (Part I, Part II, Part III and Part IV).

Continuing on the walkthrough, in this part we take the data from sessions.csv that we left aside initially and add it to the transformed and expanded data from Part IV. This part will cover, in brief, all the steps in Parts II – IV.

Data Science: A Kaggle Walkthrough – Data Transformation and Feature Extraction

March 27, 2016 / Brett Romero / 12 Comments

This article on data transformation and feature extraction is Part IV in a series looking at data science and machine learning by walking through a Kaggle competition. If you have not done so already, you are strongly encouraged to go back and read Part I, Part II and Part III.

Continuing on the walkthrough, in this part we focus on getting the data we cleaned in Part III ready for use in the classification algorithm. These steps are often referred to as data transformation and feature extraction.

Data Science: A Kaggle Walkthrough – Cleaning Data

March 14, 2016 / Brett Romero / 3 Comments

This article on cleaning data is Part III in a series looking at data science and machine learning by walking through a Kaggle competition. If you have not done so already, it is recommended that you go back and read Part I and Part II.

In this part we will focus on cleaning the data provided for the Airbnb Kaggle competition.

Category: Data Science (Page 2 of 3)

Pandas: Filtering and segmenting

Pandas: Basic data interrogation

Pandas: Reading in JSON data

Pandas: Reading in tabular data

Why the ‘boring’ part of Data Science is actually the most interesting

Forget SQL or NoSQL – 5 scenarios where you may not need a database at all

Data Science: A Kaggle Walkthrough – Creating a Model

Data Science: A Kaggle Walkthrough – Adding New Data

Data Science: A Kaggle Walkthrough – Data Transformation and Feature Extraction

Data Science: A Kaggle Walkthrough – Cleaning Data

Archives

Categories

Archives

Categories

Tags