Python Regular expression

Source: Internet
Author: User

I use both hands to achieve your dream

Python Regular Expressions
^Match start $ matches the end of the line, matching any single character other than the line break, using-The m option allows it to match a newline character as well [...] to match any of the characters in parentheses (or meaning) [^...] Match a single character or multiple characters not in parentheses*match 0 or more of the preceding expressions+match 1 or more occurrences of the preceding expression? Matches 0 or 1 occurrences of the preceding expression {n} exactly matches the number of expressions preceding the previous occurrence {n,m} matches at least n times to M times a|B matches A or B*? + ,??, {m,n}? This is *,+.,? , {m,n} becomes a non-greedy mode (RE) group regular expression and matches the text in a timely fashion (? IMX) temporarily toggles the options on the I,m or X-quake expression, if only the region is affected by the (?: RE) group regular expression and matches the remembered text (?#....) Notes(?=re) specifies the mode location to use, without a range (?! RE) uses the specified mode to take the inverse position, without a range (?<n1>..) Match \d numbers in a list [0-9] Digit \d non-digital= = [^0-9]or[^\d] \s white space character \s non-whitespace character \w alphanumeric underline word \w non-alphanumeric underline

A regular expression is a special sequence of characters that can help you check if a string matches a pattern.

Re module

The RE module uses Python to have all regular expression functionality

Re. I (re. IGNORECASE): Ignore case (full notation in parentheses)  re. M (MULTILINE): (Multiline mode, change "^", "$" behavior)  re. S (Dotall): (Point any match mode, change "." Behavior)  
Re.complit

The compile function generates a regular expression object based on a pattern string and an optional flag parameter. The object has a series of methods for regular expression matching and substitution
Format: Re.match (pattern,string,flags=0) #pattern: Regular model, string: strings to match FALGS: match pattern

A = Re.complit (R"\d*"= A.match ("ABCde")
Re.match

The Re.match function attempts to match a pattern from the location of the string, and if the match is not successful, match () returns none

Print (Re.match ('com','comwww.runcomoob'). Group ()) Print (Re.match ('com','comwww.runcomoob', re. I). Group ()) execution result: comcom
Re.seach

Re.search (pattern,string,flags=0)
The Re.search function finds a pattern match within the string, as long as the first match is found and then returns, none if the string does not match

Print (Re.search ('\dcom','www.4comrunoob.5com'). Group ()) Execution Result: 4com

* Note: match and search once matched successfully, is a match object object, and the match object object has the following methods:

1 to the included group number, usually groups () does not require parameters, returns a tuple, and the tuples in the tuple are the groups defined in the regular expression.
Importre a="123abc456"    Print(Re.search ("([0-9]*) ([a-z]*) ([ 0-9]*)", a). Group (0))#123abc456, return to the whole    Print(Re.search ("([0-9]*) ([a-z]*) ([ 0-9]*)", a). Group (1))#123    Print(Re.search ("([0-9]*) ([a-z]*) ([ 0-9]*)", a). Group (2))#ABC    Print(Re.search ("([0-9]*) ([a-z]*) ([ 0-9]*)", a). Group (3))#456

# # #group (1) lists the first bracket matching section, Group (2) lists the second Bracket matching section, and Group (3) lists the third Bracket matching section. ###

Re.findall

Re.findall traversal matches, you can get all the matching strings in the string and return a list format:
Re.findall (pattern,string,flags=0)

p = Re.compile (r'\d+')      Print(P.findall ('O1n2m3k4') ) to perform the following if: ['1','2','3','4']      ImportRe TT="Tina is a good girl, she's cool, clever, and so on ..."RR= Re.compile (r'\w*oo\w*')      Print(Rr.findall (TT))Print(Re.findall (R'(\w) *oo (\w)'TT)) The results of the execution are as follows ['Good','Cool']      [('g','D'),('C','L')]
Re.finditer

Finditer ()
Searches for a string that returns an iterator that accesses each matching result (match object) sequentially. Find the re matching so substring and put them back last night with an iterator
Format: Re.finditer (pattern,string,flags=0)

ITER = Re.finditer (r'\d+','drumm44ers Drumming, 11. Ten..')       forIinchITER:Print(i)Print(I.group ())Print(I.span ()) The results of the implementation are as follows:<_sre. Sre_match object; span= (0, 2), match=' A'> 12(0,2)      <_sre. Sre_match object; Span= (8, ten), match=' -'> 44      (8, 10)      <_sre. Sre_match object; Span= (+), match=' One'> 11      (24, 26)      <_sre. Sre_match object; Span= (+), match='Ten'> 10      (31, 33)
Re.split

Split ()
Install a string that matches strings to return a list
You can use Re.split to split strings, such as: Re.split (R ' \s+ ', text), and divide strings into a single word list by space
Format: Re.split (Pattern,string[,maxsplit])

Print (Re.split ('\d+','one1two2three3four4five5')) execution results are as follows: [ ' one ','both','three ' four','five']
Re.sub

Sub ()
Returns the replaced string after replacing each of the matched substrings in a string with re
Format: Re.sub (Pattern,repl,string,count)

ImportRetext="Jgood is a handsome boy, he's cool, clever, and so on ..."Print(Re.sub (R'\s+','-', text)) The execution results are as follows: Jgood- is-a-handsome-boy,-he- is-cool,-clever,- and-so-On ... Where the second function is a replacement string; In this case, the'-'The fourth parameter refers to the number of replacements. The default is 0, which means that each match is replaced. 

SUBN ()
Returns the number of replacements
Format:
SUBN (pattern,repl,string,count=0,flags=0)

Print(Re.subn ('[1-2]','A','123456abcdef'))Print(Re.sub ("g.t"," have",'I get A, I got B, I gut C'))Print(Re.subn ("g.t"," have",'I get A, I got B, I gut C') The results of the implementation are as follows: ('Aa3456abcdef', 2) I have A, I has B, I have C ('I have A, I has B, I have C', 3)
Difference

1. The difference between Re.match and Re.search and Re.findall:
Re.match matches only the beginning of the string, if the string does not begin to conform to the regular expression, the match fails and the function returns none;
and Re.search matches the entire string until a match is found

A=re.search ('[\d]',"Abc33"). Group ()Print(a) P=re.match ('[\d]',"Abc33")      Print(p) b=re.findall ('[\d]',"Abc33")      Print(b)
Execution Result:3none['3"3"

2. Greedy match and non-greedy match
*?,+?,??, {m,n}? In front of the *,+, and so on are greedy matches, that is, match as much as possible, after adding the number to make it an inert match

A = Re.findall (r"A (\d+?)",'a23b')      Print(a) b= Re.findall (r"A (\d+)",'a23b')      Print(b) Results of implementation: ['2']      [' at']

3. The small pits encountered with flags
Print (Re.split (' A ', ' 1a1a2a3 ', re. I)) #输出结果并未能区分大小写
This is because Re.split (pattern,string,maxsplit,flags) defaults to four parameters, and when we pass in three parameters, the system defaults to re. I is the third parameter, so it doesn't work. If you want to get here the re. I worked, written flags=re. I can.

Python Regular expression

Contact Us

The content source of this page is from Internet, which doesn't represent Alibaba Cloud's opinion; products and services mentioned on that page don't have any relationship with Alibaba Cloud. If the content of the page makes you feel confusing, please write us an email, we will handle the problem within 5 days after receiving your email.

If you find any instances of plagiarism from the community, please send an email to: info-contact@alibabacloud.com and provide relevant evidence. A staff member will contact you within 5 working days.

A Free Trial That Lets You Build Big!

Start building with 50+ products and up to 12 months usage for Elastic Compute Service

  • Sales Support

    1 on 1 presale consultation

  • After-Sales Support

    24/7 Technical Support 6 Free Tickets per Quarter Faster Response

  • Alibaba Cloud offers highly flexible support services tailored to meet your exact needs.