Regular Expressions
Say the rules I already know you are dizzy, now let us first look at some practical applications. Online test Tool http://tool.chinaz.com/regex/
The first thing you need to know is that when it comes to the regular, it's only related to strings. In the tool I have given you, every word you enter is a string.
Second, if a value in a position does not change, then no rules are required.
For example you want to use "1" to match "1", or "2" to Match "2", directly can match on. This can be done easily with Python's string operations.
Then we'll consider more of the range of characters that can appear in the same position.
Character group: [Character Group] the various characters that may appear in the same position make up a group of characters, and in regular expressions the characters are divided into classes, such as numbers, letters, punctuation, and so on. If you now ask for a position " only one number can appear ", then the character in this position can only be one of the 10 numbers, 0, 1, 2...9.
Regular |
Characters to match |
The Results |
Description |
[0123456789] |
8 |
True |
Enumerates all the valid characters in a group of characters, Any one character and the "to match" character are considered to match |
[0123456789] |
A |
False |
Cannot match because there is no "a" character in the character group |
[0-9] |
7 |
True |
can also be used-to denote a range, [0-9] and [0123456789] is a meaning |
[A-z] |
S |
True |
Similarly, if you want to match all lowercase letters, use [A-z] directly You can say |
[A-z] |
B |
True |
[A-z] means all uppercase letters |
[0-9a-fa-f] |
E |
True |
Can match numbers, case-a~f, to verify Hexadecimal characters |
Character:
Metacharacters |
Match content |
| . |
Match any character other than line break |
| \w |
Match letters or numbers or underscores |
| \s |
Match any of the whitespace characters |
| \d |
Match numbers |
| \ n |
Match a line break |
| \ t |
Match a tab |
| \b |
Match the end of a word |
| ^ |
Match the start of a string |
| $ |
Match the end of a string |
| \w |
Match a non-letter or number or underscore |
| \d |
Match non-numeric |
| \s |
Match non-whitespace characters |
| A|b |
Match character A or character B |
| () |
Matches an expression within parentheses, and also represents a group |
| [...] |
Match characters in a character group |
| [^...] |
Match all characters except characters in a character group |
Quantifiers:
Quantifiers |
Usage Notes |
| * |
Repeat 0 or more times |
| + |
Repeat one or more times |
| ? |
Repeat 0 or one time |
| N |
Repeat n times |
| {N,} |
Repeat N or more times |
| {N,m} |
Repeat N to M times |
. ^ $
| Regular |
Characters to match |
The
Results |
Description |
| Sea. |
Haiyan Hai Jiao Haidong |
Haiyan Hai Jiao Haidong |
Match all "sea." The character |
| ^ The sea. |
Haiyan Hai Jiao Haidong |
Petrels |
Only match "sea" from the beginning. |
| Sea. $ |
Haiyan Hai Jiao Haidong |
Haidong |
Match only the end of "sea. $" |
* + ? { }
| Regular |
Characters to match |
The
Results |
Description |
| Li.? |
Li Jie and Buddy and Lee two sticks |
Lijie
Li Lian
Lee |
? means repeat 0 or one time, that is, match "Li" only after An arbitrary character |
| Lee. * |
Li Jie and Buddy and Lee two sticks |
Li Jie and Buddy and Lee two sticks |
* means repeat 0 or more times, that is, match "Li" after 0 or more An arbitrary character |
| Li. + |
Li Jie and Buddy and Lee two sticks |
Li Jie and Buddy and Lee two sticks |
+ means repeat one or more times, which matches only "Li" after 1 One or more An arbitrary character |
| Li. {to} |
Li Jie and Buddy and Lee two sticks |
Li Jie and
Buddy
Li two Sticks |
{1 to 2 occurrences of any character |
Note: The previous *,+,?, etc. are greedy matches, that is, match as much as possible, followed by the number to make it an inert match
| Regular |
Characters to match |
The
Results |
Description |
| Lee. *? |
Li Jie and Buddy and Lee two sticks |
Li
Li
Li |
Lazy Matching |
Character Set [][^]
| Regular |
Characters to match |
The
Results |
Description |
| Lee [Jackie Ying two stick]* |
Li Jie and Buddy and Lee two sticks |
Lijie
Buddy
Lee Two sticks |
A character that matches the "Lee" word [Jackie two stick] any time |
| Lee [^ and]* |
Li Jie and Buddy and Lee two sticks |
Lijie
Buddy
Lee Two sticks |
Matches a character that is not "and" any time |
| [\d] |
456bdha3 |
4
5
6
3 |
Matches any number to match 4 results |
| [\d]+ |
456bdha3 |
55W
3 |
Matches any number and matches to 2 results |
Group () and or |[^]
The ID number is a 15 or 18 character string, and if it is 15 bits all??? Number, the first cannot be 0, if 18 bits, the first 17 digits are all numbers,
The bottom may be a number or x, and below we try to use the regular to indicate:
| Regular |
Characters to match |
The
Results |
Description |
| ^[1-9]\d{13,16}[0-9x]$ |
110101198001017032 |
110101198001017032 |
Indicates that a correct ID number can be matched |
| ^[1-9]\d{13,16}[0-9x]$ |
1101011980010170 |
1101011980010170 |
Indicates that the number can be matched, but this is not a correct The ID number, it's a 16-digit number |
| ^[1-9]\D{14} (\d{2}[0-9x])? $ |
1101011980010170 |
False |
The wrong ID number is now not matched () for grouping, \D{2}[0-9X] into a group, you can constrain him as a whole. The number of occurrences is 0-1 times |
| ^ ([1-9]\d{16}[0-9x]| [1-9]\d{14}) $ |
110105199812067023 |
110105199812067023 |
First match [1-9]\d{16}[0-9x] if there is no match on Just match [1-9]\d{14} |
Escape character \
In regular expressions, there are a lot of special meaning is meta-characters, such as \d and \s, if you want to match the normal "\d" instead of "number" will need to "\" to escape, become ' \ \ '.
In Python, the regular expression, or the content to be matched, is in the form of a string, which has a special meaning in the string and that itself needs to be escaped. So if the match "\d", the string to write ' \\d ', then the regular will be written in "\\\\d", so it is too troublesome. At this point we use the concept of R ' \d ', and the regular is R ' \\d '.
| Regular |
Characters to match |
The
Results |
Description |
| \d |
\d |
False |
Because \ is a character with special meaning in the regular expression, to match the \d itself, the expression \d cannot match |
| \\d |
\d |
True |
Escape \ then change to \ \ to match |
| "\\\\d" |
' \\d ' |
True |
If in Python, the ' \ ' in the string also needs to be escaped, so each string ' \ ' needs to be escaped again |
| R ' \\d ' |
R ' \d ' |
True |
Add r before the string to make the entire string not escape |
Greedy match
Greedy match: Matches the string as long as possible when matching matches, by default, greedy match
| Regular |
Characters to match |
The
Results |
Description |
| <.*> |
<script>...<script> |
<script>...<script> |
The default is greedy match mode, which matches as long as possible string |
| <.*?> |
R ' \d ' |
<script>
<script> |
Plus? To convert the greedy match pattern to a non-greedy match pattern, match the shortest possible string |
Several commonly used non-greedy matching pattern
*Repeat any number of times, but with as few repetitions as possible + repeat1 or more times, but with as few repetitions as possible? Repeat 0 or 1 times, but repeat {n,m} as little as possible. Repeat N to M, but repeat {n,} as little as possible, but repeat more than n times.
Usage of. *?
. Is any character * to take 0 to infinity length? non-greedy mode. Where together is to take as little as possible any character, generally not so alone, he mostly used in:. *? x is the length of the preceding character, until an X appears
Python re module