Splitting a dataset

栏目: IT技术 · 发布时间: 4年前

内容简介：To train any machine learning model irrespective what type of dataset is being used you have to split the dataset into training data and testing data. So, let us look into how it can be done?Here I am going to use the iris dataset and split it using the ‘t

Image by author

To train any machine learning model irrespective what type of dataset is being used you have to split the dataset into training data and testing data. So, let us look into how it can be done?

Here I am going to use the iris dataset and split it using the ‘train_test_split’ library from sklearn

from sklearn.model_selection import train_test_splitfrom sklearn.datasets import load_iris

Then I load the iris dataset into a variable.

iris = load_iris()

Which I then use to store the data and target value into two separate variables.

x, y = iris.data, iris.target

Here I have used the ‘train_test_split’ to split the data in 80:20 ratio i.e. 80% of the data will be used for training the model while 20% will be used for testing the model that is built out of it.

x_train,x_test,y_train,y_test=train_test_split(x,y,test_size=0.2,random_state=123)

As you can see here I have passed the following parameters in ‘train_test_split’:

x and y that we had previously defined
test_size: This is set 0.2 thus defining the test size will be 20% of the dataset
random_state: it controls the shuffling applied to the data before applying the split. Setting random_state a fixed value will guarantee that the same sequence of random numbers are generated each time you run the code.

When splitting a dataset there are two competing concerns:

-If you have less training data, your parameter estimates have greater variance.

-And if you have less testing data, your performance statistic will have greater variance.

The data should be divided in such a way that neither of them is too high, which is more dependent on the amount of data you have. If your data is too small then no split will give you satisfactory variance so you will have to do cross-validation but if your data is huge then it doesn’t really matter whether you choose an 80:20 split or a 90:10 split (indeed you may choose to use less training data as otherwise, it might be more computationally intensive).

以上就是本文的全部内容，希望对大家的学习有所帮助，也希望大家多多支持码农网

查看所有标签

猜你喜欢:

Splitting a dataset

本站部分资源来源于网络，本站转载出于传递更多信息之目的，版权归原作者或者来源机构所有，如转载稿涉及版权问题，请联系我们。

码农书籍

屏幕上的聪明决策

Shlomo Benartzi、Jonah Lehrer / 石磊 / 北京联合出版公司 / 2017-3 / 56.90

 为什么在手机上购物的人，常常高估商品的价值？  为什么利用网络订餐，人们更容易选择热量高的食物？  为什么网站上明明提供了所有选项，人们却还是选不到最佳的方案？  屏幕正在改变我们的思考方式，让我们变得更冲动，更容易根据直觉做出反应，进而做出错误的决策。在《屏幕上的聪明决策》一书中，什洛莫·贝纳茨教授通过引人入胜的实验及案例，揭示了究竟是什么影响了我们在屏幕上的决策。 ......一起来看看《屏幕上的聪明决策》这本书的介绍吧!

码农工具

Splitting a dataset

屏幕上的聪明决策

JSON 在线解析

UNIX 时间戳转换

HEX HSV 转换工具