Python2 爬蟲初學筆記

來源:互聯網
上載者:User

標籤:簡單   color   little   pytho   object   dex   ons   import   eth   

  

 

爬蟲,個人理解就是:利用類比“操作瀏覽器”的過程,自動擷取我們想要的資料(或者說資訊,比片啊)

為何要學爬蟲:爬取資料,為我所用(相當於可以把一類資料整合起來)

一.簡單靜態網頁爬蟲架構:

  1.Background Knowledge:URL(統一資源定位器,能協助我們定位到網頁在網路中的位置,URI 是統一資源標誌符),HTTP協議

  2.構架:

  需要一個爬蟲調度器管理下面的程式,涉及多線程管理等(比如說申請網頁的阻塞時間可以用來建立新的申請,這些資源分派由作業系統完成)

  URL管理器,防止URL重複使用,擷取URL,未爬取和已爬取的管理  

 

  

  3.工作流程:

  4.URL管理器實現方式:

    a.儲存在記憶體(set)

    b.關聯式資料庫(可永久儲存)

    c.快取資料庫(大部分公司使用這種方式)

  5.網頁下載器:

    以HTML形式儲存網頁,可以使用urllib和urllib2實現下載

    實現方法:

    a.簡單的使用urllib2.open(url)

    b.添加Request方法,發送包頭,偽裝成瀏覽器

    c.添加cookiejar cookie 容器

  

 1 # coding=utf-8 2 import urllib2 3 import cookielib 4 url = "http://www.baidu.com" 5 print ‘方法1‘ 6 #請確保url 的合法性 7 response1 = urllib2.urlopen(url) 8 if response1.getcode()==200: 9     print ‘ 讀取網頁成功‘10     print ‘ Length:‘,11     print len(response1.read())12 else:13     print ‘ 讀取網頁失敗‘14 15 print ‘Method2:‘16 request = urllib2.Request(url)17 request.add_header("usr_agent","Mozilla/6.0")18 response2 = urllib2.urlopen(request)19 if response2.getcode()==200:20     print ‘ 讀取網頁成功‘21     print ‘ Length:‘,22     print len(response2.read())23 else:24     print ‘ 讀取網頁失敗‘25 26 print ‘Method3:‘27 cj = cookielib.CookieJar()28 opener = urllib2.build_opener(urllib2.HTTPCookieProcessor(cj))29 urllib2.install_opener(opener)30 response3 = urllib2.urlopen(url)31 if response3.getcode()==200:32     print ‘ 讀取網頁成功‘33     print ‘ Length:‘,34     print len(response3.read())35     print cj36     print response3.read()37 else:38     print ‘ 讀取網頁失敗‘
View Code

  6.網頁解析器:

  以下載好的HTML當成字串,尋找出

  1.Regex匹配

  2.html.parser

   3.lxml解析器

  4.BeautifulSoup

   以DOM(Document Object Model) 結構化解析,下面是其文法

  

 1 # coding=utf-8 2 import re 3  4 from bs4 import BeautifulSoup 5 html_doc = """ 6 <html><head><title>The Dormouse‘s story</title></head> 7 <body> 8 <p class="title"><b>The Dormouse‘s story</b></p> 9 10 <p class="story">Once upon a time there were three little sisters; and their names were11 <a href="http://example.com/elsie" class="sister" id="link1">Elsie</a>,12 <a href="http://example.com/lacied" class="sister" id="link2">Lacie</a> and13 <a href="http://example.com/tillie" class="sister" id="link3">Tillie</a>;14 and they lived at the bottom of a well.</p>15 16 <p class="story">...</p>17 """18 #建立19 ccsSoup = BeautifulSoup(html_doc,‘html.parser‘,from_encoding=‘utf8‘)20 #擷取所有連結21 links= ccsSoup.find_all(‘a‘)22 for link in links:23     print link.name,link[‘href‘],link.get_text()24 print ccsSoup.p(‘class‘)25 26 print ‘正則匹配‘27 link_node = ccsSoup.find(‘a‘,href= re.compile(r"h"),class_=‘sister‘)28 print link_node29 link_node = ccsSoup.find(‘a‘,href= re.compile(r"d"))30 print link_node

  5.發送器

 

 

 

參考:  

    http://www.imooc.com/video/10686

    https://www.crummy.com/software/BeautifulSoup/bs4/doc/index.zh.html

    Regex:

      http://www.cnblogs.com/huxi/archive/2010/07/04/1771073.html

    PyCharm:使用教程

    http://blog.csdn.net/pipisorry/article/details/39909057

Python2 爬蟲初學筆記

聯繫我們

該頁面正文內容均來源於網絡整理,並不代表阿里雲官方的觀點,該頁面所提到的產品和服務也與阿里云無關,如果該頁面內容對您造成了困擾,歡迎寫郵件給我們,收到郵件我們將在5個工作日內處理。

如果您發現本社區中有涉嫌抄襲的內容,歡迎發送郵件至: info-contact@alibabacloud.com 進行舉報並提供相關證據,工作人員會在 5 個工作天內聯絡您,一經查實,本站將立刻刪除涉嫌侵權內容。

A Free Trial That Lets You Build Big!

Start building with 50+ products and up to 12 months usage for Elastic Compute Service

  • Sales Support

    1 on 1 presale consultation

  • After-Sales Support

    24/7 Technical Support 6 Free Tickets per Quarter Faster Response

  • Alibaba Cloud offers highly flexible support services tailored to meet your exact needs.