標籤:
因為業務需要,所以會有一些爬蟲的設計需求。
目前這一部分的內容都是外包項目,領導說需要根據實際情況,研究一下自己研發的可能性。
但是絕大部分這些OTA網站都做了大量的非同步載入,並且介面都做了加密處理。
也就是說,我們從控制台上面攔截下來的請求資料都是加密的資料,只有經過頁面用戶端解析後的資料,才能真正的渲染成我們想要的頁面。
因此,為了能夠正確地抓取到這些資料,我們必須使用瀏覽器核心來實現。
現在常用的瀏覽器核心工具有:
QtWebkit spynner
selenium 但是由於selenium需要開啟本地瀏覽器支援,不適合大規模爬去。
phantomjs (webkit) http://phantomjs.org
slimerjs (gecko) https://slimerjs.org
除了selenium以外,都做了測試,QtWebkit 和 slimerjs 可以成功地渲染頁面。
而phantomjs則一直卡在非同步載入無法完成的狀態。
slimerjs的代碼可以使用和phantomjs完全一樣的代碼,因為介面是完全相容的。
function render(page){page.render(‘qunar.png‘);phantom.exit();}var page = require(‘webpage‘).create();var system = require(‘system‘);page.onResourceRequested = function (request) {// 可以過濾一些不必要的請求 // system.stderr.writeLine(‘= onResourceRequested()‘); // system.stderr.writeLine(‘ request: ‘ + JSON.stringify(request, undefined, 4));};page.open(‘*ar.com/city/guangzhou/#fromDate=2015-10-01&toDate=2015-10-02&fom=qunarindex‘, function(status) {var title = page.evaluate(function() {return document.title;});console.log(‘Page title is ‘ + title);console.log("Status: " + status); //頁面返回狀態if(status === "success") {// setTimeout(render,1000,page);console.log("I am ready!");} });var t = 3;var interval = setInterval(function(){ if ( t > 0 ) { console.log(t--); } else { var htmlContent = page.evaluate(function () { return document.documentElement.outerHTML; }); console.log(htmlContent); page.render("qunar.png"); phantom.exit(); }}, 1000);
目前這個是原生的phantomjs瞎支援的代碼,通過等待一定時間,來完成頁面的ajax資源載入。
接下來是基於casperjs 的解析方案
最後,仍然要解決的問題還有不少。
phantomjs這種引擎資料持久化的問題。
webkit核心無法產生頁面的原因,甚至在測試攜 程網的適合,直接就是請求失敗。
var casper = require(‘casper‘).create({verbose: true,ogLevel: ‘debug‘, //~~添加debug參數~~userAgent: ‘Mozilla/5.0 (Macintosh; Intel Mac OS X 10_10_4) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/45.0.2454.93 Safari/537.36‘,});//監聽 requested 訊息casper.on(‘resource.requested‘, function(requestData, request){// 通過正則過濾掉一些不必要的載入if (requestData.url.match(/google|gstatic|doubleclick/)){request.abort();return;}else{this.echo(requestData.url);}});casper.start(‘*nar.com/city/taipei/‘, function() {this.echo("strat -2before");// this.scrollToBottom();//滾動到頁面底部 this.waitForSelector(‘.hotel_price‘, function() { //等到‘.tweet-row‘選取器匹配的元素出現時再執行回呼函數 this.captureSelector(‘qunar.png‘, ‘html‘); //成功時調用的函數,給整個頁面 }, function() { this.capture(‘qunar.png‘); this.die(‘Timeout reached. Fail whale?‘).exit(); //失敗時調用的函數,輸出一個訊息,並退出 }, 5000); //逾時時間,兩秒鐘後指定的選取器還沒出現,就算失敗 });casper.run();
目前而言,使用slimerjs 確實可以爬取到想要的資料內容。
使用瀏覽器核心爬取OTA資料