Here I still recommend the second method, the Find method in the current tag parse tree (currently in this HTML code block) to find a sub-tree matching the criteria and return. The Find method provides a variety of query methods, including the use of loved Regex Oh ~ after the detailed introduction.
html = soup.contents[0] #
2, with contents[], parent, nextSibling, previoussibling look for father and son brother tag
For a more convenient and flexible parsing of HTML code blocks, BeautifulSoup provides several simple ways to get directly to the parent-child sibling of the current tag block.
Suppose we have obtained the body of this tag block, we want to look for
# BODY = soup.bodyhtml = body.parent # HTML is the body of the Father Head = body.previoussibling # Head and body on the same layer, is the body of the previous brother P1 = Body.conten Ts[0] # P1, p2 is the son of the body, we use contents[0] to obtain P1P2 = p1.nextsibling # P2 and P1 on the same layer, is the P1 of the latter brother, of course body.content[1] can also get print P1.text # u ' this is Paragraphone. ' Print p2.text# u ' this is paragraphtwo. ' Note: 1, the text of each tag includes it and the text of its descendants. 2, all text has been automatically turned # for Unicode, if necessary, can be self-transcoding encode (XXX)
However, what if we are looking for ancestors or grandchildren tag?? With a while loop? No, BeautifulSoup has provided the method.
3, with Find, Findparent, findnextsibling, findprevioussibling search for ancestors or descendants tag:
With the base above, it should be well understood, such as the Find method (which I understand is the same as Findchild), that is, starting with the current node, traversing the entire subtree, and returning after finding it.
And the plural form of these methods, will find all the matching requirements of the tag, put back in the form of a list. Their correspondence is: Find->findall, findparent->findparents, findnextsibling->findnextsiblings ...
Such as:
Print Soup.findall (' P ') # [<p id= "Firstpara" align= "center" >this is paragraph <b>one</b>.</p> <p id= "Secondpara" align= "blah" >this is paragraph <b>two</b>.</p>]
Here we focus on several uses of find, other analogies:
Find (Name=none, attrs={}, Recursive=true, Text=none, **kwargs)
(PS: Only a few uses, complete please see the official link:http://www.crummy.com/software/beautifulsoup/bs3/documentation.zh.html#the%20basic% 20find%20method:%20findall%28name,%20attrs,%20recursive,%20text,%20limit,%20**kwargs%29)
1) Search tag:
Find (tagname) # searches directly for tag named tagname such as: Find (' head ') find (list) # Search for tag in list, such as: find ([' head ', ' body ']) find (d ICT) # Search for tags in dict such as: Find ({' head ': true, ' body ': true}) Find (Re.compile (')) # Search for regular tags, such as: Find (Re.compile (' ^p ') Search for a tagfind (lambda) # search function that begins with p to return a tag with a True result, such as: Find (lambda name:if len (name) = = 1) search for a length of 1 tagfind (True) # Search All Tags
2) Search properties (attrs):
Find (id= ' xxx ') # Look for the id attribute for XXX's find (attrs={id=re.compile (' xxx '), algin= ' xxx '}) # Look for the id attribute to match the regular and The Algin property is XXX's find (Attrs={id=true, algin=none}) # Looking for an id attribute but no algin attribute
3) search text (text):
Note that the search for text results in other search-giving values such as: Tag, attrs are invalidated. method is consistent with search tag
4) Recursive, limit:
Recursive=false indicates that only the immediate son is searched, otherwise the entire subtree is searched, and the default is true.
When using FindAll or a method similar to returning a list, the Limit property is used to limit the number of returns, such as FindAll (' P ', limit=2): Returns the first two tags found
4, use Next,previous to find the context tag (less)
Here we mainly look at next, next is to get the current tag of the next (in order of code from top to bottom) tag block. This is not the same as contents, do not confuse the oh ^ ^
Let's take a look at Li Zilai.
<a> a <b>b</b> <c>c</c> </a>
Let's look at the actual effect of next:
A = Soup.ab = SOUP.BN1 = B.NEXTN2 = N1.next
Output a bit:
Print a.next# u ' a ' Print n1# u ' B ' Print n2# <c>c</c>
So next is simply to get the "next" tag on the document, regardless of the location in the parse tree.
Of course there are findnext and Findallnext methods.
As for previous, which represents the last tag block, just analogy ~^ ^
Original Bo: http://www.cnblogs.com/twinsclover/archive/2012/04/26/2471704.html
moreSoupy can also easily parse the HTML document, which is based on the beautiful Soup Python library, compared with the faster to write complex HTML query,
GitHub
Https://github.com/ChrisBeaumont/soupy?utm_content=buffer6e5b7&utm_medium=social&utm_source= Twitter.com&utm_campaign=buffer
Api:
Http://soupy.readthedocs.org/en/stable/getting_started.html
Familiarity with jquery can also be used to parse HTML documents pyquery.
Python parsing html