python 簡單爬蟲

來源:互聯網
上載者:User

標籤:out   rom   res   port   skin   findall   擷取網頁   開啟   結果   

使用urllib.request 和re 模組
1 from urllib.request import * 2 import re #處理網路訪問 3 #擷取網頁 4 url = ‘https://image.baidu.com/search/index?tn=baiduimage&ct=201326592&lm=-1&cl=2&ie=gbk&word=%C3%C0%C5%AE%CD%BC%C6%AC&fr=ala&ala=1&alatpl=adress&pos=0&hs=2&xthttps=111111‘ 5 #開啟網頁 6 hmtl = urlopen(url) 7 #擷取html代碼 ,decode 解碼 8 obj = hmtl.read().decode() 9 #使用re,找出所有的objURL連結 .*?匹配所有結果10 urls = re.findall(r‘"objURL":"(.*?)"‘,obj)11 index = 112 for url in urls:13 try:14 if re.search(‘.jpg$‘,url):15 print(‘downloading........%d‘%index)16 urlretrieve(url,‘pic‘ +str(index)+ ‘.jpg‘)17 else:18 print(‘downloading........%d‘ % index)19 urlretrieve(url, ‘pic‘ + str(index) + ‘.png‘)20 index += 121 22 except Exception:23 print(‘download error....%d‘%index)24 else:25 print(‘download complete‘)

 

爬取一張圖片

使用requests 模組
1 import requests2 image_url = ‘http://www.cnblogs.com/Images/Skins/BJ2008.jpg‘3 response = requests.get(image_url)4 with open(‘outlook.jpg‘,‘wb‘) as f:5 f.write(response.content)

 

python 簡單爬蟲

聯繫我們

該頁面正文內容均來源於網絡整理,並不代表阿里雲官方的觀點,該頁面所提到的產品和服務也與阿里云無關,如果該頁面內容對您造成了困擾,歡迎寫郵件給我們,收到郵件我們將在5個工作日內處理。

如果您發現本社區中有涉嫌抄襲的內容,歡迎發送郵件至: info-contact@alibabacloud.com 進行舉報並提供相關證據,工作人員會在 5 個工作天內聯絡您,一經查實,本站將立刻刪除涉嫌侵權內容。

A Free Trial That Lets You Build Big!

Start building with 50+ products and up to 12 months usage for Elastic Compute Service

  • Sales Support

    1 on 1 presale consultation

  • After-Sales Support

    24/7 Technical Support 6 Free Tickets per Quarter Faster Response

  • Alibaba Cloud offers highly flexible support services tailored to meet your exact needs.