python爬蟲headers設定後無效的解決方案,pythonheaders
此次遇到的是一個函數使用不熟練造成的問題,但有了分析工具後可以很快定位到問題(此處推薦一個非常棒的抓包工具fiddler)
本文如下:
在爬取某個app資料時(app上的資料都是由http請求的),用Fidder分析了請求資訊,並把python的request header資訊寫在程式中進行請求資料
代碼如下
import requestsurl = 'http://xxx?startDate=2017-10-19&endDate=2017-10-19&pageIndex=1&limit=50&sort=datetime&order=desc'headers={ "Host":"xxx.com", "Connection": "keep-alive", "Accept": "application/json, text/javascript, */*; q=0.01", "User-Agent": "Mozilla/5.0 (Windows NT 6.1; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/29.0.1547.59 Safari/537.36", "X-Requested-With": "XMLHttpRequest", "Referer": "http://app.jg.eastmoney.com/html_Report/index.html", "Accept-Encoding": "gzip,deflate", "Accept-Language": "en-us,en", "Cookie":"xxx"}r = requests.get(url,headers)print (r.text)
請求成功但是,返回的是
{"Id":"6202c187-2fad-46e8-b4c6-b72ac8de0142","ReturnMsg":"載入失敗!"}
就是被發現不是正常請求被攔截了
然後我去Fidder中看剛才python發送請求的記錄 #蓋掉的兩個部分分別是Host和URL。
然後查看請求詳細資料的時候,要求標頭並沒有載入進去,User-Agent就寫著python-requests ! #要求標頭裡的UA資訊是java,python程式,有點反爬蟲意識的網站、app都會攔截掉
Header詳細資料如下
GET http://xxx?istartDate=2017-10-19&endDate=2017-10-19&pageIndex=1&limit=50&sort=datetime&order=desc &Host=xxx.com &Connection=keep-alive &Accept=application%2Fjson%2C+text%2Fjavascript%2C+%2A%2F%2A%3B+q%3D0.01 &User-Agent=Mozilla%2F5.0+%28Windows+NT+6.1%3B+WOW64%29+AppleWebKit%2F537.36+%28KHTML%2C+like+Gecko%29+Chrome%2F29.0.1547.59+Safari%2F537.36 &X-Requested-With=XMLHttpRequest &Referer=xxx &Accept-Encoding=gzip%2Cdeflate &Accept-Language=en-us%2Cen &Cookie=xxxHTTP/1.1Host: xxx.comUser-Agent: python-requests/2.18.4Accept-Encoding: gzip, deflateAccept: */*Connection: keep-aliveHTTP/1.1 200 OKServer: nginx/1.2.2Date: Sat, 21 Oct 2017 06:07:21 GMTContent-Type: application/json; charset=utf-8Content-Length: 75Connection: keep-aliveCache-Control: privateX-AspNetMvc-Version: 5.2X-AspNet-Version: 4.0.30319X-Powered-By: ASP.NET
一開始還沒發現,等我把請求的URL資訊全部讀完,才發現程式把我的要求標頭資訊當做參數放到了URL裡
那就是我請求的時候request函數Header資訊參數用錯了
又重新看了一下Requests庫的Headers參數使用方法,發現有一行代碼寫錯了,在使用request.get()方法時要把參數 “headers =“寫出來
更改如下:
import requestsurl = 'http://xxx?startDate=2017-10-19&endDate=2017-10-19&pageIndex=1&limit=50&sort=datetime&order=desc'headers={ "Host":"xxx.com", "Connection": "keep-alive", "Accept": "application/json, text/javascript, */*; q=0.01", "User-Agent": "Mozilla/5.0 (Windows NT 6.1; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/29.0.1547.59 Safari/537.36", "X-Requested-With": "XMLHttpRequest", "Referer": "http://app.jg.eastmoney.com/html_Report/index.html", "Accept-Encoding": "gzip,deflate", "Accept-Language": "en-us,en", "Cookie":"xxx"}r = requests.get(url,headers=headers)
然後去查看Fiddler中的請求。
此次python中的要求標頭已經正常了,請求詳細資料如下
GET http://xxx?startDate=2017-10-19&endDate=2017-10-19&pageIndex=1&limit=50&sort=datetime&order=desc HTTP/1.1User-Agent: Mozilla/5.0 (Windows NT 6.1; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/29.0.1547.59 Safari/537.36Accept-Encoding: gzip,deflateAccept: application/json, text/javascript, */*; q=0.01Connection: keep-aliveHost: xxx.comX-Requested-With: XMLHttpRequestReferer: http://xxxAccept-Language: en-us,enCookie: xxxHTTP/1.1 200 OKServer: nginx/1.2.2Date: Sat, 21 Oct 2017 06:42:21 GMTContent-Type: application/json; charset=utf-8Content-Length: 75Connection: keep-aliveCache-Control: privateX-AspNetMvc-Version: 5.2X-AspNet-Version: 4.0.30319X-Powered-By: ASP.NET
然後又用python程式請求了一次,結果請求成功,返回的還是
{"Id":"6202c187-2fad-46e8-b4c6-b72ac8de0142","ReturnMsg":"載入失敗!"}
因為一般cookie都會在短時間內到期,所以更新了cookie,然後請求成功
需要注意的是用程式爬蟲一定要把Header設定好,這個app如果反爬的時候封ip的話可能就麻煩了。
以上就是本文的全部內容,希望對大家的學習有所協助,也希望大家多多支援幫客之家。