標籤:機器學習 評價指標 非均衡分類
通常情況下,我們直接使用分類結果的錯誤率就可以做為該分類器的評判標準了,但是當在分類器訓練時正例數目和反例數目不相等時,這種評價標準就會出現問題。這種現象也稱為非均衡分類問題。此時有以下幾個衡量標準。
(1) 正確率<precise>和召回率<Recall>
如所示:其中準確率指預測的真實正例占所有真實正例的比例,等於TP/(TP+FP),而召回率指預測的真實正例占所有真實正例的比例,等於TP/(TP+FN)。通常我們可以很容易的構照一個高正確率或高召回率的分類器,但是很難同時保證兩者成立。如果任何樣本都被判為了正例,那麼召回率達到百分之百而此時準確率很低。構建一個同時使正確率和召回率最大的分類器是具有挑戰性的。此時我們可以用F-Score =precise*recall/(precise+ recall) 這個量來衡量,越大越好。
(2) ROC曲線
def plotROC(predStrengths, classLabels): import matplotlib.pyplot as plt cur = (1.0,1.0) #cursor ySum = 0.0 #variable to calculate AUC numPosClas = sum(array(classLabels)==1.0) yStep = 1/float(numPosClas); xStep = 1/float(len(classLabels)-numPosClas) sortedIndicies = predStrengths.argsort()#get sorted index, it's reverse fig = plt.figure() #這三行代碼用於構建畫筆 fig.clf() ax = plt.subplot(111) #loop through all the values, drawing a line segment at each point for index in sortedIndicies.tolist()[0]: if classLabels[index] == 1.0: delX = 0; delY = yStep; else: delX = xStep; delY = 0; ySum += cur[1] #draw line from cur to (cur[0]-delX,cur[1]-delY) ax.plot([cur[0],cur[0]-delX],[cur[1],cur[1]-delY], c='b') cur = (cur[0]-delX,cur[1]-delY) ax.plot([0,1],[0,1],'b--') plt.xlabel('False positive rate'); plt.ylabel('True positive rate') plt.title('ROC curve for AdaBoost horse colic detection system') ax.axis([0,1,0,1]) plt.show() print "the Area Under the Curve is: ",ySum*xStep
小村長 出處:http://blog.csdn.net/lu597203933 歡迎轉載或分享,但請務必聲明文章出處。 (新浪微博:小村長zack, 歡迎交流!)