Python re module

Source: Internet
Author: User

Regular Expressions

Say the rules I already know you are dizzy, now let us first look at some practical applications. Online test Tool http://tool.chinaz.com/regex/

The first thing you need to know is that when it comes to the regular, it's only related to strings. In the tool I have given you, every word you enter is a string.
Second, if a value in a position does not change, then no rules are required.
For example you want to use "1" to match "1", or "2" to Match "2", directly can match on. This can be done easily with Python's string operations.
Then we'll consider more of the range of characters that can appear in the same position.
Character group: [Character Group] the various characters that may appear in the same position make up a group of characters, and in regular expressions the characters are divided into classes, such as numbers, letters, punctuation, and so on. If you now ask for a position " only one number can appear ", then the character in this position can only be one of the 10 numbers, 0, 1, 2...9. 
Regular
Characters to match
The
Results
Description
[0123456789]
8
True
Enumerates all the valid characters in a group of characters,
Any one character and the "to match" character are considered to match
[0123456789]
A
False
Cannot match because there is no "a" character in the character group
[0-9]
7
True
can also be used-to denote a range, [0-9] and [0123456789] is
a meaning
[A-z]
S
True
Similarly, if you want to match all lowercase letters, use [A-z] directly
You can say
[A-z]
B
True
[A-z] means all uppercase letters
[0-9a-fa-f]
E
True
Can match numbers, case-a~f, to verify
Hexadecimal characters

Character:

Metacharacters
Match content
. Match any character other than line break
\w Match letters or numbers or underscores
\s Match any of the whitespace characters
\d Match numbers
\ n Match a line break
\ t Match a tab
\b Match the end of a word
^ Match the start of a string
$ Match the end of a string
\w
Match a non-letter or number or underscore
\d
Match non-numeric
\s
Match non-whitespace characters
A|b
Match character A or character B
()
Matches an expression within parentheses, and also represents a group
[...]
Match characters in a character group
[^...]
Match all characters except characters in a character group

Quantifiers:

Quantifiers
Usage Notes
* Repeat 0 or more times
+ Repeat one or more times
? Repeat 0 or one time
N Repeat n times
{N,} Repeat N or more times
{N,m} Repeat N to M times

. ^ $
Regular Characters to match The
Results
Description
Sea. Haiyan Hai Jiao Haidong Haiyan Hai Jiao Haidong Match all "sea." The character
^ The sea. Haiyan Hai Jiao Haidong Petrels Only match "sea" from the beginning.
Sea. $ Haiyan Hai Jiao Haidong Haidong Match only the end of "sea. $"

* + ? { }
Regular Characters to match The
Results
Description
Li.? Li Jie and Buddy and Lee two sticks

Lijie
Li Lian
Lee

? means repeat 0 or one time, that is, match "Li" only after
An arbitrary character
Lee. * Li Jie and Buddy and Lee two sticks Li Jie and Buddy and Lee two sticks
* means repeat 0 or more times, that is, match "Li" after 0 or more
An arbitrary character
Li. + Li Jie and Buddy and Lee two sticks Li Jie and Buddy and Lee two sticks
+ means repeat one or more times, which matches only "Li" after 1
One or more
An arbitrary character
Li. {to} Li Jie and Buddy and Lee two sticks

Li Jie and
Buddy
Li two Sticks

{1 to 2 occurrences of any character

Note: The previous *,+,?, etc. are greedy matches, that is, match as much as possible, followed by the number to make it an inert match

Regular Characters to match The
Results
Description
Lee. *? Li Jie and Buddy and Lee two sticks Li
Li
Li
Lazy Matching

Character Set [][^]
Regular Characters to match The
Results
Description
Lee [Jackie Ying two stick]* Li Jie and Buddy and Lee two sticks

Lijie
Buddy
Lee Two sticks

A character that matches the "Lee" word [Jackie two stick] any time
Lee [^ and]* Li Jie and Buddy and Lee two sticks

Lijie
Buddy
Lee Two sticks

Matches a character that is not "and" any time
[\d] 456bdha3

4
5
6
3

Matches any number to match 4 results
[\d]+ 456bdha3

55W
3

Matches any number and matches to 2 results

Group () and or |[^]

The ID number is a 15 or 18 character string, and if it is 15 bits all??? Number, the first cannot be 0, if 18 bits, the first 17 digits are all numbers,

The bottom may be a number or x, and below we try to use the regular to indicate:

Regular Characters to match The
Results
Description
^[1-9]\d{13,16}[0-9x]$ 110101198001017032

110101198001017032

Indicates that a correct ID number can be matched
^[1-9]\d{13,16}[0-9x]$ 1101011980010170

1101011980010170

Indicates that the number can be matched, but this is not a correct
The ID number, it's a 16-digit number
^[1-9]\D{14} (\d{2}[0-9x])? $ 1101011980010170

False

The wrong ID number is now not matched () for grouping,
\D{2}[0-9X] into a group, you can constrain him as a whole.
The number of occurrences is 0-1 times
^ ([1-9]\d{16}[0-9x]| [1-9]\d{14}) $ 110105199812067023

110105199812067023

First match [1-9]\d{16}[0-9x] if there is no match on
Just match [1-9]\d{14}

Escape character \

In regular expressions, there are a lot of special meaning is meta-characters, such as \d and \s, if you want to match the normal "\d" instead of "number" will need to "\" to escape, become ' \ \ '.

In Python, the regular expression, or the content to be matched, is in the form of a string, which has a special meaning in the string and that itself needs to be escaped. So if the match "\d", the string to write ' \\d ', then the regular will be written in "\\\\d", so it is too troublesome. At this point we use the concept of R ' \d ', and the regular is R ' \\d '.

Regular Characters to match The
Results
Description
\d \d False
Because \ is a character with special meaning in the regular expression, to match the \d itself, the expression \d cannot match
\\d \d True
Escape \ then change to \ \ to match
"\\\\d" ' \\d ' True
If in Python, the ' \ ' in the string also needs to be escaped, so each string ' \ ' needs to be escaped again
R ' \\d ' R ' \d ' True
Add r before the string to make the entire string not escape

Greedy match

Greedy match: Matches the string as long as possible when matching matches, by default, greedy match

Regular Characters to match The
Results
Description
<.*>

<script>...<script>

<script>...<script>
The default is greedy match mode, which matches as long as possible string
<.*?> R ' \d '

<script>
<script>

Plus? To convert the greedy match pattern to a non-greedy match pattern, match the shortest possible string
Several commonly used non-greedy matching pattern
*Repeat any number of times, but with as few repetitions as possible + repeat1 or more times, but with as few repetitions as possible? Repeat 0 or 1 times, but repeat {n,m} as little as possible. Repeat N to M, but repeat {n,}  as little as possible, but repeat more than n times.
Usage of. *?
. Is any character * to take 0 to infinity length? non-greedy mode. Where together is to take as little as possible any character, generally not so alone, he mostly used in:. *? x is the length of the preceding character, until an X appears 

Python re module

Contact Us

The content source of this page is from Internet, which doesn't represent Alibaba Cloud's opinion; products and services mentioned on that page don't have any relationship with Alibaba Cloud. If the content of the page makes you feel confusing, please write us an email, we will handle the problem within 5 days after receiving your email.

If you find any instances of plagiarism from the community, please send an email to: info-contact@alibabacloud.com and provide relevant evidence. A staff member will contact you within 5 working days.

A Free Trial That Lets You Build Big!

Start building with 50+ products and up to 12 months usage for Elastic Compute Service

  • Sales Support

    1 on 1 presale consultation

  • After-Sales Support

    24/7 Technical Support 6 Free Tickets per Quarter Faster Response

  • Alibaba Cloud offers highly flexible support services tailored to meet your exact needs.