Python計算KS值並繪製KS曲線

來源:互聯網
上載者:User

標籤:python   KS   大資料   

更多風控建模、大資料分析等內容請關注公眾號《大Alibaba Antifraud Service的一點一滴》

python實現KS曲線,相關使用方法請參考上篇部落格-R語言實現KS曲線

代碼如下:

####################### PlotKS ##########################def PlotKS(preds, labels, n, asc):    # preds is score: asc=1    # preds is prob: asc=0    pred = preds  # 預測值    bad = labels  # 取1為bad, 0為good    ksds = DataFrame({‘bad‘: bad, ‘pred‘: pred})    ksds[‘good‘] = 1 - ksds.bad    if asc == 1:        ksds1 = ksds.sort_values(by=[‘pred‘, ‘bad‘], ascending=[True, True])    elif asc == 0:        ksds1 = ksds.sort_values(by=[‘pred‘, ‘bad‘], ascending=[False, True])    ksds1.index = range(len(ksds1.pred))    ksds1[‘cumsum_good1‘] = 1.0*ksds1.good.cumsum()/sum(ksds1.good)    ksds1[‘cumsum_bad1‘] = 1.0*ksds1.bad.cumsum()/sum(ksds1.bad)    if asc == 1:        ksds2 = ksds.sort_values(by=[‘pred‘, ‘bad‘], ascending=[True, False])    elif asc == 0:        ksds2 = ksds.sort_values(by=[‘pred‘, ‘bad‘], ascending=[False, False])    ksds2.index = range(len(ksds2.pred))    ksds2[‘cumsum_good2‘] = 1.0*ksds2.good.cumsum()/sum(ksds2.good)    ksds2[‘cumsum_bad2‘] = 1.0*ksds2.bad.cumsum()/sum(ksds2.bad)    # ksds1 ksds2 -> average    ksds = ksds1[[‘cumsum_good1‘, ‘cumsum_bad1‘]]    ksds[‘cumsum_good2‘] = ksds2[‘cumsum_good2‘]    ksds[‘cumsum_bad2‘] = ksds2[‘cumsum_bad2‘]    ksds[‘cumsum_good‘] = (ksds[‘cumsum_good1‘] + ksds[‘cumsum_good2‘])/2    ksds[‘cumsum_bad‘] = (ksds[‘cumsum_bad1‘] + ksds[‘cumsum_bad2‘])/2    # ks    ksds[‘ks‘] = ksds[‘cumsum_bad‘] - ksds[‘cumsum_good‘]    ksds[‘tile0‘] = range(1, len(ksds.ks) + 1)    ksds[‘tile‘] = 1.0*ksds[‘tile0‘]/len(ksds[‘tile0‘])    qe = list(np.arange(0, 1, 1.0/n))    qe.append(1)    qe = qe[1:]    ks_index = Series(ksds.index)    ks_index = ks_index.quantile(q = qe)    ks_index = np.ceil(ks_index).astype(int)    ks_index = list(ks_index)    ksds = ksds.loc[ks_index]    ksds = ksds[[‘tile‘, ‘cumsum_good‘, ‘cumsum_bad‘, ‘ks‘]]    ksds0 = np.array([[0, 0, 0, 0]])    ksds = np.concatenate([ksds0, ksds], axis=0)    ksds = DataFrame(ksds, columns=[‘tile‘, ‘cumsum_good‘, ‘cumsum_bad‘, ‘ks‘])    ks_value = ksds.ks.max()    ks_pop = ksds.tile[ksds.ks.idxmax()]    print (‘ks_value is ‘ + str(np.round(ks_value, 4)) + ‘ at pop = ‘ + str(np.round(ks_pop, 4)))    # chart    plt.plot(ksds.tile, ksds.cumsum_good, label=‘cum_good‘,                         color=‘blue‘, linestyle=‘-‘, linewidth=2)    plt.plot(ksds.tile, ksds.cumsum_bad, label=‘cum_bad‘,                        color=‘red‘, linestyle=‘-‘, linewidth=2)    plt.plot(ksds.tile, ksds.ks, label=‘ks‘,                   color=‘green‘, linestyle=‘-‘, linewidth=2)    plt.axvline(ks_pop, color=‘gray‘, linestyle=‘--‘)    plt.axhline(ks_value, color=‘green‘, linestyle=‘--‘)    plt.axhline(ksds.loc[ksds.ks.idxmax(), ‘cumsum_good‘], color=‘blue‘, linestyle=‘--‘)    plt.axhline(ksds.loc[ksds.ks.idxmax(),‘cumsum_bad‘], color=‘red‘, linestyle=‘--‘)    plt.title(‘KS=%s ‘ %np.round(ks_value, 4) +                  ‘at Pop=%s‘ %np.round(ks_pop, 4), fontsize=15)    return ksds####################### over ##########################

作圖效果如下:

Python計算KS值並繪製KS曲線

聯繫我們

該頁面正文內容均來源於網絡整理,並不代表阿里雲官方的觀點,該頁面所提到的產品和服務也與阿里云無關,如果該頁面內容對您造成了困擾,歡迎寫郵件給我們,收到郵件我們將在5個工作日內處理。

如果您發現本社區中有涉嫌抄襲的內容,歡迎發送郵件至: info-contact@alibabacloud.com 進行舉報並提供相關證據,工作人員會在 5 個工作天內聯絡您,一經查實,本站將立刻刪除涉嫌侵權內容。

A Free Trial That Lets You Build Big!

Start building with 50+ products and up to 12 months usage for Elastic Compute Service

  • Sales Support

    1 on 1 presale consultation

  • After-Sales Support

    24/7 Technical Support 6 Free Tickets per Quarter Faster Response

  • Alibaba Cloud offers highly flexible support services tailored to meet your exact needs.