1. 用php擷取遠程網址header頭資訊的方法,這在採集時很有用,他可以讓你判斷出來,遠程檔案或網頁是否正常,是否是404頁.
$url = 'http://www.example.com';
print_r(get_headers($url));
上例的輸出類似於:
Array
(
[0] => HTTP/1.1 200 OK
[1] => Date: Sat, 29 May 2004 12:28:13 GMT
[2] => Server: Apache/1.3.27 (Unix) (Red-Hat/Linux)
[3] => Last-Modified: Wed, 08 Jan 2003 23:11:55 GMT
[4] => ETag: "3f80f-1b6-3e1cb03b"
[5] => Accept-Ranges: bytes
[6] => Content-Length: 438
[7] => Connection: close
[8] => Content-Type: text/html
)
---------------------------------------------------
$url = 'http://www.example.com';
print_r(get_headers($url, 1));
Array
(
[0] => HTTP/1.1 200 OK
[Date] => Sat, 29 May 2004 12:28:14 GMT
[Server] => Apache/1.3.27 (Unix) (Red-Hat/Linux)
[Last-Modified] => Wed, 08 Jan 2003 23:11:55 GMT
[ETag] => "3f80f-1b6-3e1cb03b"
[Accept-Ranges] => bytes
[Content-Length] => 438
[Connection] => close
[Content-Type] => text/html
)
get_headers
是用來取得遠程伺服器的回應標頭資訊的.用返回的第一個數組再加上正則就可以判斷遠程地址是否為200正常網頁
--------------------------------------------------------------
2. 用curl
CURLOPT_NOBODY參數只抓取header頭資訊
curl函數真是個好東西,curl參數裡有一項可以配置只抓取遠程網頁的header頭資訊
如下代碼,他指定了curl抓的內容中包含header頭,並且不要body內容.
function get_header($url){
$ch = curl_init();
curl_setopt($ch, CURLOPT_URL, $url);
curl_setopt($ch, CURLOPT_HEADER, true);
curl_setopt($ch, CURLOPT_NOBODY,true);
curl_setopt($ch, CURLOPT_RETURNTRANSFER,true);
curl_setopt($ch, CURLOPT_FOLLOWLOCATION,true);
curl_setopt($ch, CURLOPT_AUTOREFERER,true);
curl_setopt($ch, CURLOPT_TIMEOUT,30);
curl_setopt($ch, CURLOPT_HTTPHEADER, array(
'Accept: */*',
'User-Agent: Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1; SV1)',
'Connection: Keep-Alive'));
$header = curl_exec($ch);
return $header;
}