I use both hands to achieve your dream
Python Regular Expressions
^Match start $ matches the end of the line, matching any single character other than the line break, using-The m option allows it to match a newline character as well [...] to match any of the characters in parentheses (or meaning) [^...] Match a single character or multiple characters not in parentheses*match 0 or more of the preceding expressions+match 1 or more occurrences of the preceding expression? Matches 0 or 1 occurrences of the preceding expression {n} exactly matches the number of expressions preceding the previous occurrence {n,m} matches at least n times to M times a|B matches A or B*? + ,??, {m,n}? This is *,+.,? , {m,n} becomes a non-greedy mode (RE) group regular expression and matches the text in a timely fashion (? IMX) temporarily toggles the options on the I,m or X-quake expression, if only the region is affected by the (?: RE) group regular expression and matches the remembered text (?#....) Notes(?=re) specifies the mode location to use, without a range (?! RE) uses the specified mode to take the inverse position, without a range (?<n1>..) Match \d numbers in a list [0-9] Digit \d non-digital= = [^0-9]or[^\d] \s white space character \s non-whitespace character \w alphanumeric underline word \w non-alphanumeric underline
A regular expression is a special sequence of characters that can help you check if a string matches a pattern.
Re module
The RE module uses Python to have all regular expression functionality
Re. I (re. IGNORECASE): Ignore case (full notation in parentheses) re. M (MULTILINE): (Multiline mode, change "^", "$" behavior) re. S (Dotall): (Point any match mode, change "." Behavior)
Re.complit
The compile function generates a regular expression object based on a pattern string and an optional flag parameter. The object has a series of methods for regular expression matching and substitution
Format: Re.match (pattern,string,flags=0) #pattern: Regular model, string: strings to match FALGS: match pattern
A = Re.complit (R"\d*"= A.match ("ABCde")
Re.match
The Re.match function attempts to match a pattern from the location of the string, and if the match is not successful, match () returns none
Print (Re.match ('com','comwww.runcomoob'). Group ()) Print (Re.match ('com','comwww.runcomoob', re. I). Group ()) execution result: comcom
Re.seach
Re.search (pattern,string,flags=0)
The Re.search function finds a pattern match within the string, as long as the first match is found and then returns, none if the string does not match
Print (Re.search ('\dcom','www.4comrunoob.5com'). Group ()) Execution Result: 4com
* Note: match and search once matched successfully, is a match object object, and the match object object has the following methods:
1 to the included group number, usually groups () does not require parameters, returns a tuple, and the tuples in the tuple are the groups defined in the regular expression.
Importre a="123abc456" Print(Re.search ("([0-9]*) ([a-z]*) ([ 0-9]*)", a). Group (0))#123abc456, return to the whole Print(Re.search ("([0-9]*) ([a-z]*) ([ 0-9]*)", a). Group (1))#123 Print(Re.search ("([0-9]*) ([a-z]*) ([ 0-9]*)", a). Group (2))#ABC Print(Re.search ("([0-9]*) ([a-z]*) ([ 0-9]*)", a). Group (3))#456
# # #group (1) lists the first bracket matching section, Group (2) lists the second Bracket matching section, and Group (3) lists the third Bracket matching section. ###
Re.findall
Re.findall traversal matches, you can get all the matching strings in the string and return a list format:
Re.findall (pattern,string,flags=0)
p = Re.compile (r'\d+') Print(P.findall ('O1n2m3k4') ) to perform the following if: ['1','2','3','4'] ImportRe TT="Tina is a good girl, she's cool, clever, and so on ..."RR= Re.compile (r'\w*oo\w*') Print(Rr.findall (TT))Print(Re.findall (R'(\w) *oo (\w)'TT)) The results of the execution are as follows ['Good','Cool'] [('g','D'),('C','L')]Re.finditer
Finditer ()
Searches for a string that returns an iterator that accesses each matching result (match object) sequentially. Find the re matching so substring and put them back last night with an iterator
Format: Re.finditer (pattern,string,flags=0)
ITER = Re.finditer (r'\d+','drumm44ers Drumming, 11. Ten..') forIinchITER:Print(i)Print(I.group ())Print(I.span ()) The results of the implementation are as follows:<_sre. Sre_match object; span= (0, 2), match=' A'> 12(0,2) <_sre. Sre_match object; Span= (8, ten), match=' -'> 44 (8, 10) <_sre. Sre_match object; Span= (+), match=' One'> 11 (24, 26) <_sre. Sre_match object; Span= (+), match='Ten'> 10 (31, 33)
Re.split
Split ()
Install a string that matches strings to return a list
You can use Re.split to split strings, such as: Re.split (R ' \s+ ', text), and divide strings into a single word list by space
Format: Re.split (Pattern,string[,maxsplit])
Print (Re.split ('\d+','one1two2three3four4five5')) execution results are as follows: [ ' one ','both','three ' four','five']
Re.sub
Sub ()
Returns the replaced string after replacing each of the matched substrings in a string with re
Format: Re.sub (Pattern,repl,string,count)
ImportRetext="Jgood is a handsome boy, he's cool, clever, and so on ..."Print(Re.sub (R'\s+','-', text)) The execution results are as follows: Jgood- is-a-handsome-boy,-he- is-cool,-clever,- and-so-On ... Where the second function is a replacement string; In this case, the'-'The fourth parameter refers to the number of replacements. The default is 0, which means that each match is replaced.
SUBN ()
Returns the number of replacements
Format:
SUBN (pattern,repl,string,count=0,flags=0)
Print(Re.subn ('[1-2]','A','123456abcdef'))Print(Re.sub ("g.t"," have",'I get A, I got B, I gut C'))Print(Re.subn ("g.t"," have",'I get A, I got B, I gut C') The results of the implementation are as follows: ('Aa3456abcdef', 2) I have A, I has B, I have C ('I have A, I has B, I have C', 3)Difference
1. The difference between Re.match and Re.search and Re.findall:
Re.match matches only the beginning of the string, if the string does not begin to conform to the regular expression, the match fails and the function returns none;
and Re.search matches the entire string until a match is found
A=re.search ('[\d]',"Abc33"). Group ()Print(a) P=re.match ('[\d]',"Abc33") Print(p) b=re.findall ('[\d]',"Abc33") Print(b)
Execution Result:3none['3"3"
2. Greedy match and non-greedy match
*?,+?,??, {m,n}? In front of the *,+, and so on are greedy matches, that is, match as much as possible, after adding the number to make it an inert match
A = Re.findall (r"A (\d+?)",'a23b') Print(a) b= Re.findall (r"A (\d+)",'a23b') Print(b) Results of implementation: ['2'] [' at']
3. The small pits encountered with flags
Print (Re.split (' A ', ' 1a1a2a3 ', re. I)) #输出结果并未能区分大小写
This is because Re.split (pattern,string,maxsplit,flags) defaults to four parameters, and when we pass in three parameters, the system defaults to re. I is the third parameter, so it doesn't work. If you want to get here the re. I worked, written flags=re. I can.
Python Regular expression