python numpy 把資料向量化

來源:互聯網
上載者:User

標籤:

某公司的資料處理。

比如有資料:

0 1 20100304 51.14 5 62.2 0 0.0 1 6.82 4.08666666667 4 30 0 20100307 2.47333333333 3 75.9333333333 0 0.0 114 13.5666666667 2.86666666667 0 00 2 20100318 2.49333333333 3 58.6666666667 0 0.0 1 5.22666666667 1.61333333333 0 00 2 20100329 9.14666666667 14 70.6428571429 0 0.0 1 31.84 7.51333333333 11 70 0 20100401 3.36666666667 5 71.1 1 75.0 1 1.86 1.29333333333 0 00 1 20100401 7.44666666667 11 58.1909090909 0 0.0 51 40.8066666667 12.5066666667 2 10 0 20100401 1.0 1 138.0 0 0.0 1 1.0 1.0 0 01 0 20100428 1.0 1 85.0 0 0.0 1 0.0 0.0 0 01 0 20100429 1.0 1 128.0 0 0.0 1 24.54 3.59333333333 0 00 0 20100506 1.92 2 67.0 0 0.0 44 6.65333333333 1.91333333333 4 1

要求:

1、除去第1列和第2列。
2、對剩下的日期列提取特徵,以月份為特徵,形式:201003

3、對剩下的數字列,用最大值和最小值的區間,劃分十份,每一份就是一個特徵,形式:[1,5)、[5,9)...[21,25],注意最後一個區間閉合了。

4、用上面擷取的特徵,把未經處理資料向量化,形式:

1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1 1 0 0 0 0 0 0 0 0 0 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1 0 0 0 0 0 0 0 0 0 1 0 0 0 0 0 0 0 0 0 1 0 0 0 0 0 0 0 0 0 1 0 0 0 0 0 0 0 0 0 1 0 0 0 0 0 0 0 0 0 


開寫:

先下載安裝pyExcelerator:http://sourceforge.net/projects/pyexcelerator/,解壓後setup.py install

# -*- coding: cp936 -*-import numpy as npfrom numpy import * import pprint from pyExcelerator import *#all_col_feature存放特徵們all_col_feature = []for i in range(2,13):    #擷取[日期]列特徵:201003,201004...201408...    if i == 2 :        FILENAME = "lost5"        date = [e[0:6] for e in loadtxt(FILENAME , delimiter = " " , usecols=(2,) , dtype=str)]        date = list(set(date)) #去重        date.sort()        all_col_feature.append(date)    #擷取[數字]列特徵:取最大值,最小值,劃分10份,每一份就是一個特徵,[1,5),[5,9)..[21,25]    else:        digit_col = filter(lambda e:e!=0,loadtxt(FILENAME , delimiter = " " , usecols=(i,) , dtype=float))#擷取資料後去0        digit_col.sort()        max_num = digit_col[-1]        min_num = digit_col[0]        block_size = (max_num - min_num) /10        col_feature = []        for i in range(11):            col_feature.append(min_num+i*block_size)        all_col_feature.append(col_feature)w = Workbook() #建立一個活頁簿    ws = w.add_sheet('Zero_One_Data') #建立一個工作表row_index = 0col_index = 0for i in range(len(all_col_feature)):    if i == 0:        for j in all_col_feature[i]:            ws.write(row_index,col_index,str(j))            col_index += 1    else:        for j in range(len(all_col_feature[i])-1):            ws.write(row_index,col_index,str(all_col_feature[i][j])+"~"+str(all_col_feature[i][j+1]))            col_index += 1row_index += 1out = file('out.txt','w')res = []for row in loadtxt(FILENAME , delimiter = " " , usecols=(2,3,4,5,6,7,8,9,10,11,12) , dtype=str):    #資料整理成這個形式:every_row_feature_index = [0, 1, 1, 1, 0, 0, 0, 1, 1, 1, 1]    #比如:第一個0表示在[日期]列的特徵下標為0,-1表示NONE值    #比如:第一個1表示在第二列的特徵下標為1,心好累,我這都要注釋?    every_row_feature_index = []    a = [i for i in range(len(all_col_feature[0])) if row[0][0:6] <= all_col_feature[0][i]]    every_row_feature_index.append(a[0])        for col in range(1,11):        b = [i for i in range(len(all_col_feature[col])) if float(row[col]) < float(all_col_feature[col][i])]        if len(b) == 0:#max_num            every_row_feature_index.append(len(all_col_feature)-2)        else:            every_row_feature_index.append(b[0]-1) #-1表NONE                #產生向量    #就是把every_row_feature_index = [0, 1, 1, 1, 0, 0, 0, 1, 1, 1, 1]轉成下面的形式:    #[1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]    zero_one_every_row_res = []    col_index = 0    for i in range(len(every_row_feature_index)):        if i == 0: #處理[日期]列            for j in range(every_row_feature_index[i]):                out.write('0'+' ')                ws.write(row_index,col_index,'0')                col_index += 1            out.write('1'+' ')            ws.write(row_index,col_index,'1')            col_index += 1            for j in range(every_row_feature_index[i]+1,len(all_col_feature[i])):                out.write('0'+' ')                ws.write(row_index,col_index,'0')                col_index += 1        else:#處理其他列            if every_row_feature_index[i] == -1: #-1表NONE值 向量全為0                for j in range(len(all_col_feature[i])-1):                    out.write('0'+' ')                    ws.write(row_index,col_index,'0')                    col_index += 1            else:                for j in range(every_row_feature_index[i]):                    out.write('0'+' ')                    ws.write(row_index,col_index,'0')                    col_index += 1                out.write('1'+' ')                ws.write(row_index,col_index,'1')                col_index += 1                for j in range(every_row_feature_index[i]+1,len(all_col_feature[i])-1):                    out.write('0'+' ')                    ws.write(row_index,col_index,'0')                    col_index += 1    out.write('\n')    row_index += 1    if row_index == 100:#為了更快看到結果,只設定100行,去掉可得到全部        breakout.close()w.save('out.xls')


看看結果:



著作權聲明:本文為博主原創文章,未經博主允許不得轉載。

python numpy 把資料向量化

聯繫我們

該頁面正文內容均來源於網絡整理,並不代表阿里雲官方的觀點,該頁面所提到的產品和服務也與阿里云無關,如果該頁面內容對您造成了困擾,歡迎寫郵件給我們,收到郵件我們將在5個工作日內處理。

如果您發現本社區中有涉嫌抄襲的內容,歡迎發送郵件至: info-contact@alibabacloud.com 進行舉報並提供相關證據,工作人員會在 5 個工作天內聯絡您,一經查實,本站將立刻刪除涉嫌侵權內容。

A Free Trial That Lets You Build Big!

Start building with 50+ products and up to 12 months usage for Elastic Compute Service

  • Sales Support

    1 on 1 presale consultation

  • After-Sales Support

    24/7 Technical Support 6 Free Tickets per Quarter Faster Response

  • Alibaba Cloud offers highly flexible support services tailored to meet your exact needs.