Visualize Missing Values with Missingno

栏目: IT技术 · 发布时间: 4年前

内容简介:Explore the missing values in your dataset.Data is the new fuel. However, the raw data is cheap. We need to process it well to take the most value out of it. Complex, well-structured models are as good as the data we feed to it. Thus, data needs to be clea

Visualize Missing Values with Missingno

Explore the missing values in your dataset.

Photo by Irina on Unsplash

Data is the new fuel. However, the raw data is cheap. We need to process it well to take the most value out of it. Complex, well-structured models are as good as the data we feed to it. Thus, data needs to be cleaned and processed thoroughly in order to build robust and accurate models.

One of the issues that we are likely to encounter in raw data is missing values. Consider a case where we have features (columns in a dataframe) on some observations (rows in a dataframe). If we do not have the value in a particular row-column pair, then we have a missing value. We may have only a few missing values or half of an entire column may be missing. In some cases, we can just ignore or drop the rows or columns with missing values. On the other, there might be some cases in which we cannot afford to drop even a single missing value. In any case, handling missing values process starts with exploring them in the dataset.

Pandas provides functions to check the number of missing values in the dataset. Missingno library takes it one step further and provides the distribution of missing values in the dataset by informative visualizations. Using the plots of missingno , we are able to see where the missing values are located in each column and if there is a correlation between missing values of different columns. Before handling missing values, it is very important to explore them in the dataset. Thus, I consider missingno as a highly valuable asset in data cleaning and preprocessing steps.

In this post, we will explore the functionalities of missingno plot by going through some examples.

Let’s first try to explore a dataset about the movies on streaming platforms. The dataset is available here on kaggle.

import numpy as np
import pandas as pddf = pd.read_csv("/content/MoviesOnStreamingPlatforms.csv")
print(df.shape)
df.head()

以上就是本文的全部内容,希望本文的内容对大家的学习或者工作能带来一定的帮助,也希望大家多多支持 码农网

查看所有标签

猜你喜欢:

本站部分资源来源于网络,本站转载出于传递更多信息之目的,版权归原作者或者来源机构所有,如转载稿涉及版权问题,请联系我们

掘金大数据

掘金大数据

程新洲、朱常波、晁昆 / 机械工业出版社 / 2019-1 / 59.00元

在数据横向融合的时代,充分挖掘数据金矿及盘活数据资产,是企业发展和转型的关键所在。电信运营商以其数据特殊性,必将成为大数据领域的领航者、生力军。各行业的大数据从业者要如何从电信业的大数据中挖掘价值呢? 本书彻底揭开电信运营商数据的神秘面纱,系统介绍了大数据的发展历程,主要的数据挖掘方法,电信运营商在网络运行及业务运营方面的数据资源特征,基于用户、业务、网络、终端及内在联系的电信运营商大数据分......一起来看看 《掘金大数据》 这本书的介绍吧!

随机密码生成器
随机密码生成器

多种字符组合密码

Base64 编码/解码
Base64 编码/解码

Base64 编码/解码

RGB CMYK 转换工具
RGB CMYK 转换工具

RGB CMYK 互转工具