Python BeautifulSoup 簡單筆記

來源:互聯網
上載者:User

標籤:

 

2013-07-30 22:54 by 江湖麼名, 2359 閱讀, 0 評論, 收藏, 編輯

Beautiful Soup 是用 Python 寫的一個 HTML/XML 的解析器,它可以很好的處理不規範標記並產生剖析樹。通常用來分析爬蟲抓取的web文檔。對於 不規則的 Html文檔,也有很多的補全功能,節省了開發人員的時間和精力。

Beautiful Soup 的官方文檔齊全,將官方給出的例子實踐一遍就能掌握。官方英文文檔,中文文檔

一 安裝 Beautiful Soup

安裝 BeautifulSoup 很簡單,下載 BeautifulSoup  源碼。解壓運行

python setup.py install 即可。

測試安裝是否成功。鍵入 import BeautifulSoup 如果沒有異常,即成功安裝

二 使用 BeautifulSoup

1. 匯入BeautifulSoup ,建立BeautifulSoup 對象

from BeautifulSoup import BeautifulSoup           # HTMLfrom BeautifulSoup import BeautifulStoneSoup      # XMLimport BeautifulSoup                              # ALL                                                                                               doc = [    ‘<html><head><title>Page title</title></head>‘,    ‘<body><p id="firstpara" align="center">This is paragraph <b>one</b>.‘,    ‘<p id="secondpara" align="blah">This is paragraph <b>two</b>.‘,    ‘</html>‘]# BeautifulSoup 接受一個字串參數soup = BeautifulSoup(‘‘.join(doc))

2. BeautifulSoup對象簡介

用BeautifulSoup 解析 html文檔時,BeautifulSoup將 html文檔類似 dom文檔樹一樣處理。BeautifulSoup文檔樹有三種基本對象。

2.1. soup BeautifulSoup.BeautifulSoup

type(soup)<class ‘BeautifulSoup.BeautifulSoup‘>

2.2. 標記 BeautifulSoup.Tag

type(soup.html)<class ‘BeautifulSoup.Tag‘>

2.3 文本 BeautifulSoup.NavigableString

type(soup.title.string)<class ‘BeautifulSoup.NavigableString‘>

3. BeautifulSoup 剖析樹

3.1 BeautifulSoup.Tag對象方法

擷取 標記對象(Tag)

標記名擷取法 ,直接用 soup對象加標記名,返回 tag對象.這種方式,選取唯一標籤的時候比較有用。或者根據樹的結構去選取,一層層的選擇

>>> html = soup.html>>> html<html><head><title>Page title</title></head><body><p id="firstpara" align="center">This is paragraph<b>one</b>.</p><p id="secondpara" align="blah">This is paragraph<b>two</b>.</p></body></html>>>> type(html)<class ‘BeautifulSoup.Tag‘>>>> title = soup.title<title>Page title</title>

content方法

content方法 根據文檔樹進行搜尋,返回標記對象(tag)的列表

>>> soup.contents[<html><head><title>Page title</title></head><body><p id="firstpara" align="center">This is paragraph<b>one</b>.</p><p id="secondpara" align="blah">This is paragraph<b>two</b>.</p></body></html>]
>>> soup.contents[0].contents[<head><title>Page title</title></head>, <body><p id="firstpara" align="center">This is paragraph<b>one</b>.</p><p id="secondpara" align="blah">This is paragraph<b>two</b>.</p></body>]>>> len(soup.contents[0].contents)2>>> type(soup.contents[0].contents[1])<class ‘BeautifulSoup.Tag‘>

使用contents向後遍曆樹,使用parent向前遍曆樹

next 方法

擷取樹的子代元素,包括 Tag 對象 和 NavigableString 對象。。。

>>> head.next<title>Page title</title>>>> head.next.nextu‘Page title‘
>>> p1 = soup.p>>> p1<p id="firstpara" align="center">This is paragraph<b>one</b>.</p>>>> p1.nextu‘This is paragraph‘

nextSibling 下一個兄弟對象 包括 Tag 對象 和 NavigableString 對象

>>> head.nextSibling<body><p id="firstpara" align="center">This is paragraph<b>one</b>.</p><p id="secondpara" align="blah">This is paragraph<b>two</b>.</p></body>>>> p1.next.nextSibling<b>one</b>

與 nextSibling 相似的是 previousSibling,即上一個兄弟節點。

replacewith方法

將對象替換為,接受字串參數

>>> head = soup.head>>> head<head><title>Page title</title></head>>>> head.parent<html><head><title>Page title</title></head><body><p id="firstpara" align="center">This is paragraph<b>one</b>.</p><p id="secondpara" align="blah">This is paragraph<b>two</b>.</p></body></html>>>> head.replaceWith(‘head was replace‘)>>> head<head><title>Page title</title></head>>>> head.parent>>> soup<html>head was replace<body><p id="firstpara" align="center">This is paragraph<b>one</b>.</p><p id="secondpara" align="blah">This is paragraph<b>two</b>.</p></body></html>>>>

搜尋方法

搜尋提供了兩個方法,一個是 find,一個是findAll。這裡的兩個方法(findAll和 find)僅對Tag對象以及,頂層剖析對象有效,但 NavigableString不可用。

findAll(name, attrs, recursive, text, limit, **kwargs)

接受一個參數,標記名

尋找文檔所有 P標記,返回一個列表

>>> soup.findAll(‘p‘)[<p id="firstpara" align="center">This is paragraph<b>one</b>.</p>, <p id="secondpara" align="blah">This is paragraph<b>two</b>.</p>]>>> type(soup.findAll(‘p‘))<type ‘list‘>

尋找 id="secondpara"的 p 標記,返回一個結果集

>>> pid = type(soup.findAll(‘p‘,id=‘firstpara‘))>>> pid<class ‘BeautifulSoup.ResultSet‘>

傳一個屬性或多個屬性對

>>> p2 = soup.findAll(‘p‘,{‘align‘:‘blah‘})>>> p2[<p id="secondpara" align="blah">This is paragraph<b>two</b>.</p>]>>> type(p2)<class ‘BeautifulSoup.ResultSet‘>

利用Regex

>>> soup.findAll(id=re.compile("para$"))[<p id="firstpara" align="center">This is paragraph<b>one</b>.</p>, <p id="secondpara" align="blah">This is paragraph<b>two</b>.</p>]

讀取和修改屬性

>>> p1 = soup.p>>> p1<p id="firstpara" align="center">This is paragraph<b>one</b>.</p>>>> p1[‘id‘]u‘firstpara‘>>> p1[‘id‘] = ‘changeid‘>>> p1<p id="changeid" align="center">This is paragraph<b>one</b>.</p>>>> p1[‘class‘] = ‘new class‘>>> p1<p id="changeid" align="center" class="new class">This is paragraph<b>one</b>.</p>>>>

剖析樹基本方法就這些,還有其他一些,以及如何配合Regex。具體請看官方文檔

 

3.2 BeautifulSoup.NavigableString對象方法

NavigableString  對象方法比較簡單,擷取其內容

>>> soup.title<title>Page title</title>>>> title = soup.title.next>>> titleu‘Page title‘>>> type(title)<class ‘BeautifulSoup.NavigableString‘>>>> title.stringu‘Page title‘

至於如何遍曆樹,進而分析文檔,已經 XML 文檔的分析方法,可以參考官方文檔。

Python BeautifulSoup 簡單筆記

聯繫我們

該頁面正文內容均來源於網絡整理,並不代表阿里雲官方的觀點,該頁面所提到的產品和服務也與阿里云無關,如果該頁面內容對您造成了困擾,歡迎寫郵件給我們,收到郵件我們將在5個工作日內處理。

如果您發現本社區中有涉嫌抄襲的內容,歡迎發送郵件至: info-contact@alibabacloud.com 進行舉報並提供相關證據,工作人員會在 5 個工作天內聯絡您,一經查實,本站將立刻刪除涉嫌侵權內容。

A Free Trial That Lets You Build Big!

Start building with 50+ products and up to 12 months usage for Elastic Compute Service

  • Sales Support

    1 on 1 presale consultation

  • After-Sales Support

    24/7 Technical Support 6 Free Tickets per Quarter Faster Response

  • Alibaba Cloud offers highly flexible support services tailored to meet your exact needs.