More Performance Evaluation Metrics You Should Know for Classification Problems

栏目: IT技术 · 发布时间: 4年前

内容简介:Precisionis the ratio ofLow precision: more the number of False positives the model predicts lesser the precision.Recall (Sensitivity)is the ratio of

The equations of 4 key classification metrics

Recall versus Precision

Precisionis the ratio of True Positives to all the positives predicted by the model.

Low precision: more the number of False positives the model predicts lesser the precision.

Recall (Sensitivity)is the ratio of True Positives to all the positives in your Dataset.

Low recall: more the number of False Negatives the model predicts lesser the recall.

The idea of recall and precision seems to be abstract. Let me illustrate the difference in three real cases.
  • the result of TP will be that the COVID 19 residents diagnosed with COVID-19.
  • the result of TN will be that healthy residents are with good health.
  • the result of FP will be that those actually healthy residents are predicted as COVID 19 residents.
  • the result of FN will be that those actual COVID 19 residents are predicted as the healthy residents

In case 1, which scenario do you think will have the highest cost?

Imagine that if we predict COVID 19 residents as healthy patients and they do not need to quarantine, there would be a massive number of COVID 19 infection. The cost of f alse negative is much higher the cost of f alse positives.

  • the result of TP will be that spam emails are placed in the spam folder.
  • the result of TN will be that important emails are received.
  • the result of FP will be that important emails are placed in the spam folder.
  • the result of FN will be that spam emails are received.

In case 2, which scenario do you think will have the highest cost?

Well, since missing important emails will clearly be more of a problem than receiving spam, we can say that in this case, FP will have a higher cost than FN.

  • the result of TP will be that bad loans are correctly predicted as bad loans.
  • the result of TN will be that good loans are correctly predicted as good loans.
  • the result of FP will be that (actual) good loans are incorrectly predicted as bad loans.
  • the result of FN will be that (actual) bad loans are incorrectly predicted as good loans.

In case 3, which scenario do you think will have the highest cost?

The banks would lose a bunch amount of money if the actual bad loans are predicted as good loans due to loans not being repaid. In other hands, banks wont be able to make more revenue if the actual good loans are predicted as bad loans. Therefore, the cost of False Negatives is much higher the cost of False Positives. Imagine that


In practice, the cost of false negative is not the same as the cost of false positive depending on the different specific cases. It is evident that not only should we calculate accuracy, but we should also evaluate our model using other metrics, for example, Recall and Precision .

以上就是本文的全部内容,希望对大家的学习有所帮助,也希望大家多多支持 码农网






凯尼格 / 高巍 / 人民邮电出版社 / 2008-2-1 / 30.00元

作者以自己1985年在Bell实验室时发表的一篇论文为基础,结合自己的工作经验扩展成为这本对C程序员具有珍贵价值的经典著作。写作本书的出发点不是要批判C语言,而是要帮助C程序员绕过编程过程中的陷阱和障碍。.. 全书分为8章,分别从词法分析、语法语义、连接、库函数、预处理器、可移植性缺陷等几个方面分析了C编程中可能遇到的问题。最后,作者用一章的篇幅给出了若干具有实用价值的建议。.. 本书......一起来看看 《C陷阱与缺陷》 这本书的介绍吧!

SHA 加密
SHA 加密

SHA 加密工具

