標籤:簡單 color little pytho object dex ons import eth
爬蟲,個人理解就是:利用類比“操作瀏覽器”的過程,自動擷取我們想要的資料(或者說資訊,比片啊)
為何要學爬蟲:爬取資料,為我所用(相當於可以把一類資料整合起來)
一.簡單靜態網頁爬蟲架構:
1.Background Knowledge:URL(統一資源定位器,能協助我們定位到網頁在網路中的位置,URI 是統一資源標誌符),HTTP協議
2.構架:
需要一個爬蟲調度器管理下面的程式,涉及多線程管理等(比如說申請網頁的阻塞時間可以用來建立新的申請,這些資源分派由作業系統完成)
URL管理器,防止URL重複使用,擷取URL,未爬取和已爬取的管理
3.工作流程:
4.URL管理器實現方式:
a.儲存在記憶體(set)
b.關聯式資料庫(可永久儲存)
c.快取資料庫(大部分公司使用這種方式)
5.網頁下載器:
以HTML形式儲存網頁,可以使用urllib和urllib2實現下載
實現方法:
a.簡單的使用urllib2.open(url)
b.添加Request方法,發送包頭,偽裝成瀏覽器
c.添加cookiejar cookie 容器
1 # coding=utf-8 2 import urllib2 3 import cookielib 4 url = "http://www.baidu.com" 5 print ‘方法1‘ 6 #請確保url 的合法性 7 response1 = urllib2.urlopen(url) 8 if response1.getcode()==200: 9 print ‘ 讀取網頁成功‘10 print ‘ Length:‘,11 print len(response1.read())12 else:13 print ‘ 讀取網頁失敗‘14 15 print ‘Method2:‘16 request = urllib2.Request(url)17 request.add_header("usr_agent","Mozilla/6.0")18 response2 = urllib2.urlopen(request)19 if response2.getcode()==200:20 print ‘ 讀取網頁成功‘21 print ‘ Length:‘,22 print len(response2.read())23 else:24 print ‘ 讀取網頁失敗‘25 26 print ‘Method3:‘27 cj = cookielib.CookieJar()28 opener = urllib2.build_opener(urllib2.HTTPCookieProcessor(cj))29 urllib2.install_opener(opener)30 response3 = urllib2.urlopen(url)31 if response3.getcode()==200:32 print ‘ 讀取網頁成功‘33 print ‘ Length:‘,34 print len(response3.read())35 print cj36 print response3.read()37 else:38 print ‘ 讀取網頁失敗‘View Code
6.網頁解析器:
以下載好的HTML當成字串,尋找出
1.Regex匹配
2.html.parser
3.lxml解析器
4.BeautifulSoup
以DOM(Document Object Model) 結構化解析,下面是其文法
1 # coding=utf-8 2 import re 3 4 from bs4 import BeautifulSoup 5 html_doc = """ 6 <html><head><title>The Dormouse‘s story</title></head> 7 <body> 8 <p class="title"><b>The Dormouse‘s story</b></p> 9 10 <p class="story">Once upon a time there were three little sisters; and their names were11 <a href="http://example.com/elsie" class="sister" id="link1">Elsie</a>,12 <a href="http://example.com/lacied" class="sister" id="link2">Lacie</a> and13 <a href="http://example.com/tillie" class="sister" id="link3">Tillie</a>;14 and they lived at the bottom of a well.</p>15 16 <p class="story">...</p>17 """18 #建立19 ccsSoup = BeautifulSoup(html_doc,‘html.parser‘,from_encoding=‘utf8‘)20 #擷取所有連結21 links= ccsSoup.find_all(‘a‘)22 for link in links:23 print link.name,link[‘href‘],link.get_text()24 print ccsSoup.p(‘class‘)25 26 print ‘正則匹配‘27 link_node = ccsSoup.find(‘a‘,href= re.compile(r"h"),class_=‘sister‘)28 print link_node29 link_node = ccsSoup.find(‘a‘,href= re.compile(r"d"))30 print link_node
5.發送器
參考:
http://www.imooc.com/video/10686
https://www.crummy.com/software/BeautifulSoup/bs4/doc/index.zh.html
Regex:
http://www.cnblogs.com/huxi/archive/2010/07/04/1771073.html
PyCharm:使用教程
http://blog.csdn.net/pipisorry/article/details/39909057
Python2 爬蟲初學筆記