python爬蟲scrapy之貸聯盟黑名單爬取

來源:互聯網
上載者:User

1、建立項目

scrapy startproject ppd
2,爬取單頁,主要用xpath

spider裡面的源碼

from scrapy.spiders import Spiderfrom scrapy.selector import Selectorfrom ppd.items import BlackItemclass PpdSpider(Spider):    name = "ppd"    allowed_domains = ["dailianmeng.com"]    start_urls = [        "http://www.dailianmeng.com/p2pblacklist/index.html"    ]    def parse(self, response):      sites = response.xpath('//*[@id="yw0"]/table/tbody/tr')      items = []      for site in sites:        item = BlackItem()        item['name'] = site.xpath('td[1]/text()').extract()        item['idcard'] = site.xpath('td[2]/text()').extract()        item['mobile']=site.xpath('td[3]/text()').extract()        item['email']=site.xpath('td[4]/text()').extract()        item['total']=site.xpath('td[5]/text()').extract()        item['bepaid']=site.xpath('td[6]/text()').extract()        item['notPaid']=site.xpath('td[7]/text()').extract()        item['time']=site.xpath('td[8]/text()').extract()        item['loanAmount']=site.xpath('td[9]/text()').extract()        items.append(item)      return items

items.py增加的源碼

class BlackItem(Item):    name = Field()    idcard = Field()    mobile=Field()    email=Field()    total=Field()    bepaid=Field()    notPaid=Field()    time=Field()    loanAmount=Field()

結果成功跑出單頁結果,但是屬性排序按照屬性名稱字的大小寫字母,下一步要改成我定義的順序。

3,按照指定屬性順序輸出
因為scrapy本來按照字母順序輸出屬性和屬性值,現在我想改為按照我指定的順序:

首先在目錄中spider裡面,建立一個檔案,命名為csv_item_exporter.py

from scrapy.conf import settingsfrom scrapy.contrib.exporter import CsvItemExporterclass MyProjectCsvItemExporter(CsvItemExporter):    def __init__(self, *args, **kwargs):        delimiter = settings.get('CSV_DELIMITER', ',')        kwargs['delimiter'] = delimiter        fields_to_export = settings.get('FIELDS_TO_EXPORT', [])        if fields_to_export :            kwargs['fields_to_export'] = fields_to_export        super(MyProjectCsvItemExporter, self).__init__(*args, **kwargs)

然後,再去setting.py,加入以下代碼,其中下面屬性的順序是自己決定的,還有字典開始的檔案名稱也要替換成自己的project 名字:

FEED_EXPORTERS = {    'csv': 'ppd.spiders.csv_item_exporter.MyProjectCsvItemExporter',} #jsuser為工程名FIELDS_TO_EXPORT = [      'name',     'idcard',      'mobile',      'email',      'total',      'bepaid',      'notPaid',      'time',      'loanAmount']


結果:

name,idcard,mobile,email,total,bepaid,notPaid,time,loanAmount餘良鋒,61250119890307****,13055099***,233424611@qq.com,3000.00,1063.01,999.89,2013-10-11,3個月張栩,44152219890923****,15767638***,466780713@qq.com,3000.00,2319.84,819.56,2013-09-09,3個月孫福東,37150219890919****,15194000***,1275972787@qq.com,3000.00,2075.14,1018.55,2013-09-25,3個月李其印,45012119870211****,13481120***,993412914@qq.com,3050.00,2127.64,167.99,2013-04-08,1年吳必擁,45232819810201****,13977369***,13977369850@139.com,3000.00,2670.40,524.01,2013-06-07,6個月單長江,32072319820512****,18094220***,13851220018@136.com,8900.00,6302.04,1521.78,2013-07-22,6個月鄭其睦,35042619890215****,15959783***,271320236@qq.com,5000.00,3278.60,425.51,2013-04-08,1年吳文豪,44190019890929****,13267561***,234663601@qq.com,6000.00,579.79,463.40,2013-10-09,1年鐘華,45060319870526****,18277072***,48959434@qq.com,5700.00,3141.24,957.50,2013-08-07,6個月湯雙傑,34082119620804****,13329062***,332416587@qq.com,100000.00,105293.45,9111.54,2012-11-19,1年黃河,43240219791103****,13786520***,zhuanghe8589@126.com,6700.00,4795.24,2307.54,2013-06-21,6個月孫景昌,13092119850717****,15127714***,492101828@qq.com,3000.00, ,455.71,2013-10-18,1年高義,42050319740831****,15337410***,10855219@qq.com,3000.00, ,965.51,2013-10-17,6個月曹成均,41088119720221****,18639192***,1404816232@qq.com,3300.00,1781.64,838.18,2013-06-17,8個月張銀球,33032519761109****,13806800***,1838599723@qq.com,60000.00, ,19407.50,2013-10-16,6個月


4,爬取所有頁面的資料

主要有以下特點:(1)頁面自動擷取;(2)寫入迴圈,爬取多個頁面,(3)速度相對於selenium更快

from scrapy.spiders import Spiderfrom scrapy.selector import Selectorfrom ppd.items import BlackItemclass PpdSpider(Spider):  name = "ppd"  allowed_domains = ["dailianmeng.com"]  start_urls = []  #start_urls.append("http://www.dailianmeng.com/p2pblacklist/index.html")  #total_page = 164
  page_re=request.get('http://www.dailianmeng.com/p2pblacklist/index.html')  page_info=page_re.find_element_by_css_selector('#yw0 > div.summary')  # 第 2446-2448 條, 共 2448 條.  pages=page_info.text.split(',')[1]  pages=int(int(pages[3:6])/15)  size_page = pages
  size_page = 165  start_page = 1  for pge in range(start_page,start_page + size_page):    start_urls.append('http://www.dailianmeng.com/p2pblacklist/index.html?P2pBlacklist_page='+str(pge))  def parse(self, response):      sites = response.xpath('//*[@id="yw0"]/table/tbody/tr')      items = []      for site in sites:        item = BlackItem()        item['name'] = site.xpath('td[1]/text()').extract()        item['idcard'] = site.xpath('td[2]/text()').extract()        item['mobile']=site.xpath('td[3]/text()').extract()        item['email']=site.xpath('td[4]/text()').extract()        item['total']=site.xpath('td[5]/text()').extract()        item['bepaid']=site.xpath('td[6]/text()').extract()        item['notPaid']=site.xpath('td[7]/text()').extract()        item['time']=site.xpath('td[8]/text()').extract()        item['loanAmount']=site.xpath('td[9]/text()').extract()        items.append(item)      return items


聯繫我們

該頁面正文內容均來源於網絡整理,並不代表阿里雲官方的觀點,該頁面所提到的產品和服務也與阿里云無關,如果該頁面內容對您造成了困擾,歡迎寫郵件給我們,收到郵件我們將在5個工作日內處理。

如果您發現本社區中有涉嫌抄襲的內容,歡迎發送郵件至: info-contact@alibabacloud.com 進行舉報並提供相關證據,工作人員會在 5 個工作天內聯絡您,一經查實,本站將立刻刪除涉嫌侵權內容。

A Free Trial That Lets You Build Big!

Start building with 50+ products and up to 12 months usage for Elastic Compute Service

  • Sales Support

    1 on 1 presale consultation

  • After-Sales Support

    24/7 Technical Support 6 Free Tickets per Quarter Faster Response

  • Alibaba Cloud offers highly flexible support services tailored to meet your exact needs.