PHP Regular Expression basic syntax and php Regular Expression
Previous
Regular Expressions are a syntax rule used to describe character arrangement and matching modes. It is mainly used for string mode segmentation, matching, search and replacement operations. In PHP, regular expressions are generally a procedural description of a text pattern that consists of regular characters and special characters (similar to wildcards. Regular expressions have three functions: 1. Matching, which is often used to extract information from strings; 2. Replacing matching text with new text; 3. Split a string into a group of smaller information blocks. This article will introduce the basic regular expression syntax in PHP in detail.
[Note] the detailed information about the javascript Regular Expression is now
History
There are two sets of Regular Expression function libraries in PHP. The two functions are similar, but the execution efficiency is slightly different: one is provided by the PCRE (Perl Compatible Regular Expression) Library, a function that uses the prefix "prefix _". Another set of functions provided by the POSIX (Portable Operating System Interface of Unix) Extension uses the prefix "ereg _".
PCRE comes from the Perl language, and Perl is one of the most powerful languages for string operations. The initial version of PHP is a product developed by Perl. PCRE syntax supports more features, more powerful than POSIX syntax
POSIX was used before PHP4. Currently, mainstream PCRE is used.
A regular expression acts as a matching pattern. It is composed of atoms (common characters, such as characters a to z), special characters (metacharacters, such as *, +, and? And pattern modifier.
Delimiters
The backslash (/) is often used as the delimiter, for example, "/apple /". You only need to place the content of the pattern to be matched between delimiters. The Delimiter is not limited to "/". Any character except letters, numbers, and diagonal lines "\" can be used as delimiters, such as "#", "|", and "!". And so on.
/<\/\ W +>/-- The backslash is used as the delimiters. | (\ d {3})-\ d + | Sm -- the vertical line is used. | "is used as the delimiters! ^ (? I) php [34]! -- Use a vertical line "!" Valid as delimiters/href = '(. *)' -- invalid delimiters; end delimiters 1-\ d3-\ d3-\ d4 | -- invalid delimiters; missing start delimiters
If the delimiter needs to be matched in the mode, it must be escaped using a backslash. If separators often appear in the mode, a better choice is to use other separators to improve readability.
/http:\/\//#http://#
Metacharacters
The power of a regular expression is that it can have the ability to select and repeat in a pattern. Some characters are given special meanings, so that they do not simply represent themselves. In the pattern, such encoded characters with special meanings are called metacharacters.
There are two different metacharacters: one can be used anywhere outside the Chinese brackets of the pattern, and the other must be used in square brackets.
[1] The metacharacters used outside square brackets are as follows:
\ Is generally used to escape characters ^ the start position of the asserted target (or the first row in multi-row mode) $ the end position of the asserted target (or the end of the row in multi-row mode ). match any character except the line break (default) [start character class definition] End character class definition | start the end mark of an optional branch (the start mark of the Child group? As a quantizer, it indicates 0 or 1 match. Located behind the quantifiers, it is used to change the greedy characteristics of quantifiers * quantifiers, 0 or multiple matches + quantifiers, 1 or multiple matches {custom quantifiers start Mark} custom quantifiers end mark
[Note] outside the character class, the period in the mode matches any character in the target string, including non-printable characters, but the line break is not included by default. If PCRE_DOTALL is set, the period matches the line break.
[2] The section in the Chinese braces of the pattern is called "character class ". Only the following metacharacters are available in a character class:
\ Escape Character ^ only when it is used as the first character (in square brackets), it indicates that the character class has a reversed-flag character range.
Backslash
Backlash has multiple usage methods. First, if it is followed by a non-alphanumeric character, it indicates that the special meaning represented by this character is canceled. This method uses the backslash as an escape character, which is available both inside and outside the character class.
For example, if you want to match a "*" character, you must write it as "\ *" in the mode "\*". This applies when a character is not escaped and has a special meaning. However, it is safe to add a backslash to the front of a non-alphanumeric character when it needs to match the original text. If you want to match a backslash, use "\" in the mode "\\"
The backslash has special meanings in single quotes and double quotation marks. Therefore, to match a backslash, The backslash must be written as "\\\\" in the mode "\\\\"
[Note] "// \\/" is used as a string. The backslash will be escaped. The Escape result is //. This is the pattern obtained by the Regular Expression Engine, the Regular Expression Engine also considers "\" as an escape mark, which will escape the separator/to get an error. Therefore, four backslashes are required to match a backslash.
The second use of the backslash provides a method to control the visible encoding of non-printable characters. Except that the binary 0 ends a mode, it does not strictly limit the appearance of non-printable characters (itself). However, when a mode is edited and prepared in a text editor, it is easier to use the following escape sequence than to use binary characters.
\ A bell character (hexadecimal 07) \ cx "control-x", x is any character \ e escape (hexadecimal 1B) \ f form feed (hexadecimal 0C) \ n line feed (hexadecimal 0A) \ p {xx} A character conforming to the xx attribute \ P {xx} A character not conforming to the xx attribute \ r press enter (hexadecimal 0D) \ t horizontal tab (hexadecimal 09) \ xhh hh hexadecimal character \ ddd octal character, or backward reference
Character class
\ D any decimal number \ D any non-decimal number \ h any horizontal white space character \ H any non-horizontal white space character \ s any blank character \ S any non-blank character \ v any vertical white space character \ V any non-vertical white space character \ w any word character \ W any non-word character
[Note] A word character refers to any letter, number, or underline. That is to say, any character that can constitute a perl word
The fourth method of backlash is some simple assertions. An assertion specifies a condition that must be matched at a specific position and does not consume any characters from the target string. Backlash assertions include:
\ B word boundary \ B Non-word boundary \ A target's start position (independent from multiline mode) \ Z target's end position or line break at the end (independent from multiline Mode) \ z target end position (independent from multi-row mode) \ G first match position in target
[Note] \ A, \ Z, \ z assertions are different from the traditional ^ and $, because they always match the start and end of the target string, and are not limited by the pattern modifier.
Anchor
Outside a character class, in the default matching mode, ^ is an asserted that the current matching point is located at the beginning of the target string. Inside a character class, ^ indicates that the characters described in this character class are reversed.
^ Is not necessarily the first character of the mode, but it should be the first character of an optional branch. If all selected branches start with ^, this means that if the mode is limited to match only the beginning of the target, it is called a "Fastening" mode.
Dollar signs are used to assert that the current match point is at the end of the target string, or when the target string ends with a linefeed, the current match point is at the linefeed position (default ). $ Does not have to be the last character of the mode, but if it is in an optional branch, it should be at the end of the branch. The dollar symbol has no special meaning in the character class.
The meaning of the dollar sign can be changed to match only the end of the string by setting PCRE_DOLLAR_ENDONLY during compilation or matching. This does not affect the behavior of \ Z assertions.
^ And the meaning of dollar characters will change when the PCRE_MULTILINE option is set. In this case, they match the start and end of each line break and the first character. In addition, they also match the start and end of the target string. For example, the pattern/^ abc dollar sign/matches the target string "def \ nabc" in multiline mode, but not in normal cases. Therefore, since all the available branches start with ^, This is the fastening mode in single row mode, but in multi-row mode, this is not fastening. The PCRE_DOLLAR_ENDONLY option is invalid after PCRE_MULTILINE is set.
Dollar sign $
Character class
The left square brackets start the description of a character class and end with square brackets. A separate right brace has no special meaning. If a right square bracket needs to be a member of a character class, you can write it at the first character of the character class (if ^ is used, it is second) or use an escape character
A character class matches a single character in the target string. This character must be a combination of character sets defined in the character class unless ^ is used to reverse the character class. If ^ needs to be a member of a character class, make sure it is not the first character of the class, or escape it.
For example, the character class [aeiou] matches all lowercase vowels, while the [^ aeiou] matches all non-Vowel characters. Note: ^ is only a convenient symbol for specifying characters that do not exist in the character class through enumeration. Instead of assertion, it will still consume one character from the target string, and if the current match point is at the end of the target string, the match will fail
When case-insensitive matching is set, any character classes are both case-insensitive versions, A [aeiou] That is case insensitive matches "a" and "A" at the same time, and is case insensitive [^ aeiou] does not match "A" at the same time"
Line breaks have no special meaning in character classes and are irrelevant to the PCRE_DOTALL or PCRE_MULTILINE options. A character class, such as [^ a], always matches line breaks.
In a character class, a hyphen (minus sign-) can be used to specify the range from one character to another. For example, if [d-m] matches all characters between d and m, the set is closed. If the hyphen itself needs to be described in a character class, it must be transferred or appear in a location that is not interpreted as a range, such as the start or end position of the character class.
Right brackets cannot be used after a character range description. For example, a pattern [W-] 46] is interpreted as a character class containing W and-, followed by the string "46]", therefore, it can match "W46]" or "-46]". However, if brackets are escaped, they are interpreted as the end of the range, therefore, [W-\] 46] is interpreted as a separate character class containing all characters in the range of W to] and 4 and 6. The brackets described in octal or hexadecimal format can also be used as the end point of the range.
Range Operations are sorted in ASCII order. They can be used to specify numbers for characters, such as [\ 000-\ 037]. If a range containing letters is used in case-insensitive match mode, the format matches the uppercase and lowercase letters at the same time. For example, [W-c] is equivalent to [] [\ ^ _ 'wxyzabc] In case-insensitive matching, and if "fr" (France) is used) when the locale table is set for, [\ xc8-xcb] will match the accent E character in all modes
Character Classes \ d, \ D, \ s, \ S, \ w, and \ W can also appear in a character class, it is used to add the matched character classes to the new custom character classes. For example, [\ dABCDEF] matches any valid hexadecimal number. You can use ^ to easily define strict character classes. For example, [^ \ W _] matches any letter or number but does not match the underline.
All non-alphanumeric characters except \,-, ^ (at the starting position), and ending] are non-special characters in the character class, and no escape is harmful. The pattern Terminator is always a special character in the expression and must be escaped.
Optional path
The vertical line character (|) is used as an optional path in separation mode. For example, the mode gilbert | Sullivan matches "gilbert" or "sullivan ". The vertical bars can appear in any number of modes and allow available optional paths (matching empty strings ). Each optional path is tried from left to right for matching processing, and the first matching is successful. If the available path is in the sub-group, "successful match" indicates that the sub-mode branch and other parts of the main mode are matched at the same time.
Pattern Modifier
Pattern modifiers are used outside the regular expression delimiters, generally after the last slash. The pattern modifier can adjust the interpretation of a regular expression and extend some functions of a regular expression in operations such as matching and replacement. It can also be used in combination, enhanced the processing capability of regular expressions.
Pattern modifiers are helpful for writing concise and short expressions. Some common pattern modifiers and their function descriptions are listed below.
I. It is case-insensitive for mode matching.
M treats the string as multiple rows. The default regular start ^ and end $ take the target string as a single line of characters. If the m modifier is used, the start and end points to each row of the string.
Point character in s mode. Match All characters, including line breaks
The white space in the x mode is ignored unless it has been escaped.
E only uses the preg_replace () function to replace the reverse reference in the replacement string. Use it as the PHP code to evaluate the value and use the result to replace the searched string.
U this modifier reverses the value of the matching quantity so that it is not the default repetition, but becomes behind? And is not compatible with Perl. You can also set the U modifier in the mode or add a question mark after the digit? To enable this option
The dollar sign in D mode matches only the end of the target string. If there are no options, if the last character is a line break, the dollar sign will also match before this character. If the m modifier is set, ignore this option.
Sub-group
Child Groups (Child patterns) are defined by parentheses and can be nested. Marking part of a mode as a sub-group (sub-mode) mainly involves two tasks:
1. localize the available branches. For example, if cat (arcat | erpillar |) matches one of cat, cataract, and caterpillar, it matches "cataract", "erpillar", and empty strings.
2. Set the Sub-Group as the capture sub-group. After the entire pattern match, the part of the target string that matches the Sub-Group will be passed back to the caller through the ovector parameter of pcre_exec. The order in which left parentheses appear from left to right is the subscript of the corresponding sub-group (starting from 1). You can use these subscript numbers to obtain the capture sub-pattern matching result.
For example, if the string "the red king" is matched using the (red | white) (king | queen) mode, the result of the pattern match is array ("red king ", "red king", "red", "king"), where 0th elements are the results of the entire pattern match, the following three elements are matched by the three sub-groups in sequence. Their subscripts are 1, 2, and 3, respectively.
In fact, the two functions of parentheses are not always useful. We often need to use sub-groups for grouping, but do not need to capture them (separately. Followed by the left parenthesis defined by the Child Group followed by the string "? : "The Sub-Group will not be captured independently and will not affect the calculation of the subsequent sub-group sequence number. For example, if the string "the white queen" matches the pattern ((? : Red | white) (king | queen), the matched results will be array ("white queen", "white queen", "white queen ") and king | queen sub-groups. The maximum number of sub-groups to be captured is 99, and the maximum number of all sub-groups (including captured and non-captured) allowed to be owned is 200.
To facilitate shorthand, if you need to set options at the beginning of a non-capturing sub-group, the option letter can be located? And:, for example:
(?i:saturday|sunday)(?:(?i)saturday|sunday)
The above two methods are actually the same. Because the optional branch tries each branch from left to right, and the option is not reset before the child mode ends, and because the option setting affects other branches, all the above modes match "SUNDAY" and "Saturday"
In PHP 4.3.3, you can use (? P <name> pattern) syntax. This sub-mode will appear in the matching results in both its name and order (numerical subscript). in PHP 5.2.2, two kinds of sub-group naming syntax are added :(? <Name> pattern) and (? 'Name' pattern)
You can select a sub-group in a regular expression if multiple matches are required. In order to allow multiple sub-groups to share one backward reference number ,(? | The syntax allows you to copy numbers.
Consider the following regular expression matching Sunday:
(?:(Sat)ur|(Sun))day
Here, when backward reference 1 is null, Sun is stored in backward Reference 2. When backward Reference 2 does not exist, the Sat is stored in backward reference 1. Use (? | Modify the mode to fix this problem:
(?|(Sat)ur|(Sun))day
In this mode, Sun and Sat are both stored in backward reference 1.
Quantifiers
The number of repetitions is specified by quantifiers and can be followed by the following elements: individual characters, which can be escaped; metacharacters; character classes; Backward references; sub-group (unless it is an asserted)
A general repeated quantizer specifies the number of times a minimum value matches a maximum value. Two numbers are enclosed in curly brackets, and the two numbers are defined using the syntax separated by commas. Both values must be less than 65536, and the first number must be less than or equal to the second value.
For example, z {2, 4} matches "zz", "zzz", and "zzzz ". A single right curly braces are not special characters. If the second digit is omitted, but the comma still exists, there is no upper limit. If the second digit and the comma are ignored, the quantifiers are limited to a matching of the specified number of times. For example, [aeiou] {3,} matches at least three consecutive vowels, but can also match more, while \ d {8} can only match eight numbers. The left curly braces are considered to be a common character when they are not allowed to use quantifiers or do not match the quantifiers syntax. For example, {, 6} is not a quantizer and will match four characters "{, 6}" according to the original text }"
The quantizer {0} is authorized, and it will lead to behavior that the previous item and quantifiers do not exist.
For convenience (and historical compatibility), the three most commonly used quantifiers have single-character abbreviations.
* Equivalent to {0,} + is equivalent to {1 ,}? Equivalent to {0, 1}
A subpattern that does not match any character followed by a 0 or multiple character quantifiers can be used to construct an infinite loop without an upper limit. For example: (?) *
By default, quantifiers are "greedy", that is, they will match as many characters as possible (until the maximum number of matching times is allowed) without causing a pattern match failure ). A typical example of this problem is to try to match the C language annotations. All contents that appear between/* and */are considered comments. In the middle of the comment, separate * And/are allowed/
One attempt to match the C annotation is to use the mode /\*. * \ */. If this mode is applied to the string "/* first comment */not comment/* second comment */", it will match the wrong result, that is, the entire string, which is caused by the greedy nature of quantifiers. It will try to match as many characters as possible.
However, if a quantizer is followed by one? (Question mark) Mark, it will become a lazy (non-Greedy) pattern, it no longer matches as much as possible, but as few as possible. So the mode /\*.*? \ */The statement will be correctly executed in C's Annotation matching. The meaning of each quantizer is not changed, but is it added? Change the number of matching times. Do not set? This usage is confused with its usage as a quantizer. Because of its two usage methods, sometimes it may contain quantifiers, such as \ d ?? \ D is more inclined to match a number, but it can also accept matching of two numbers to achieve the purpose of matching the entire pattern. : In mode \ w \ d ?? \ D \ w for example, for the string "a33a", although \ d ?? It is not greedy, but the entire pattern does not match because of greedy. Therefore, the final choice is still matching a number.
If the PCRE_UNGREEDY option is set (an option that is unavailable in perl), the quantifiers are not greedy by default. However, a single quantizer can follow one? To make it greedy. In other words, the PCRE_UNGREEDY option reverses greedy default behavior.
Followed by a "+", the quantifiers are "possession. It will eat as many characters as possible and will not focus on other subsequent modes, such ,. * abc matches "aabc",. * + abc does not match because. * + will eat the entire string, resulting in the remaining pattern not matching. Since PHP 4.3.3, you can use the placeholder (+) modifier to increase the speed.
When a sub-group is modified by a quantizer with a minimum quantity greater than 1 or a maximum number of quantifiers, more storage is required for the compilation mode according to the minimum or maximum quantity.
If a mode uses. * Or. {0,} starts and the PCRE_DOTALL option is enabled (equivalent to/s of perl), that is, allow. match the line break, then the mode will be implicitly tightened, because no matter what, next we will try every character position in the target string, so after the first time, there is no retry point for all matches at any location. PCRE will try to process this mode like \. When we know that the target string does not contain a line break, it is worth setting PCRE_DOTALL to get this optimization when the mode starts with. *, or we choose to use ^ to explicitly specify the anchor.
When a sub-group is captured repeatedly, the result of the sub-group captured is the value captured in the last iteration. For example, (tweedle [dume] {3} \ s *) + matches the string "tweedledum tweedledee", and the obtained Sub-Group capture result is "tweedledee ". However, if it is a nested capture sub-group, the corresponding capture value may be set to the previous iteration. For example,/(a | (B) +/matches the string "aba". The result of the second capture sub-group is "B"
Backward reference
Outside a character class, the backslash is followed by a number greater than 0 (possibly with a single digit) and is the backward reference of a capture group that appeared before the pattern.
If the number followed by the backslash is less than 10, it is always a back reference, and if there are not so many capture groups in the mode, an error is thrown. In other words, the number of referenced parentheses cannot be less than 10.
A backward reference will directly match the content actually captured by the referenced capture group in the target string, instead of matching the content in the sub-group mode. Therefore, mode (sens | respons) e and \ 1ibility will match "sense and sensibility" and "response and responsibility" instead of "sense and responsibility"
If the backend reference is forced to perform case-sensitive matching, for example ((? I) Raah) \ s + \ 1 matches "Raah" and "Raah", but does not match "Raah", even if the original capturing Sub-Group itself is case-insensitive
More than one subgroup may be referenced to the reference. A sub-group may not be used for a specific match. In this case, any backward reference to this sub-group will fail. For example, mode (a | (bc) \ 2 always fails when it matches a string starting with "a" instead of "bc. As there may be up to 99 back-references, all numbers followed by the backslash may be a potential back-reference count. If the mode is followed by a numeric character after backward reference, some delimiters must be used to end the backward reference syntax. If the PCRE_EXTENDED option is set, use spaces. In other cases, you can use an empty comment.
If a backward reference appears in the referenced Sub-Group, its matching will fail. For example, (a \ 1) won't get any matching results. However, this type of reference can be used for repeated internal sub-patterns. For example, the pattern (a | B \ 1) + matches any number of strings consisting of "a", "aba", and "ababba. In the iteration process of each sub-mode, the Back Reference matches the string that this sub-group matches during the previous iteration. In order to do this kind of work, the pattern must meet such a condition. During the first iteration of the pattern, the pattern must be able to ensure that it does not need to match the backward reference. This condition can be implemented using an optional path like in the preceding example, or by modifying the backward reference using a quantizer with the minimum value of 0.
After PHP 5.2.2, the \ g escape sequence can be used for absolute and relative references of the sub-mode. The escape sequence must be followed by an unsigned number or a negative number. You can use parentheses to enclose the number. The sequences \ 1, \ g1, \ g {1} are synonyms. This method can eliminate the ambiguity generated when the backslash is used to closely follow the Numerical Description of the reverse reference. This escape sequence is helpful for distinguishing back-to-back and octal numeric characters, and also makes it clearer that back-to-back references are followed by an original matching number, such as \ g {2} 1
The \ g escape sequence follows a negative number to indicate a relative backward reference. For example: (foo) (bar) \ g {-1} can match the string "foobarbar", (foo) (bar) \ g {2} can match "foobarfoo ". This is an optional scheme in the long mode to keep track of the subgroup numbers referenced by the previous special subgroup
Backward reference also supports the use of sub-group name syntax description, such (? P = name) or PHP 5.2.2 can be used \ k <name> or \ k'name '. In addition, \ k {name} And \ g {name} are supported in PHP 5.2.4.
Assertions
An assertion is a test of the characters before or after the current matching position. It does not actually consume any characters. Simple asserted codes include \ B, \ B, \ A, \ Z, \ z, ^, and $. More complex assertions are encoded as child groups. It has two types: forward-looking assertions (test from the current position forward) and backward-looking assertions (test from the current position backward)
An assertion sub-group is still matched in the normal way. The difference is that it will not cause the current match point to change. The positive assertions in the forward-looking assertions (asserted that this match is true) are = ", Negative assertions start with" (?!" . For example, \ w + (? =;) Match a word followed by a semicolon, but the matching result does not contain a semicolon, foo (?! Bar) matches all "foo" strings that are not followed by "bar. Pay attention to a similar mode (?! Foo) bar, which cannot be used to find all bar matches that are not "foo" before. It will find any bar match, because (?! Foo) This assertion is always TRUE when bar is the next three characters. This is what the forward-looking assertions need to achieve.
The front assertions in the rear-view assertions use <= ", Negative assertions start with" (? <!" Start. For example ,(? <! Foo) bar is used to find any bar that is not "foo ". The content of the Post-Zhan assertions is strictly limited to match a fixed-length string only. However, if there are multiple optional branches, they do not need to have the same length. For example (? <= Bullock | donkey) is allowed, (? <! Dogs? | Cats ?) This will cause a compilation error. It is allowed to match strings of different lengths in the parent branch. Compared with perl5.005, it requires multiple branches to match strings of the same length. (? <= AB (c | de) is not allowed because its single top-level branch can match two different lengths, however, it can use two top-level branches (? <= Abc | abde). For each optional branch, the current position is temporarily moved to the fixed width before the current position to be matched. If not, the matching fails. The combination of post-Zhan assertions and one-time sub-groups can be used to match the end of a string. An example is to give the end of a string on a one-time sub-group.
Multiple assertions (in any order) can appear at the same time. For example (? <= \ D {3 })(? <! 999) foo matches the string "foo" with three numbers but not "999 ". Note that each assertion is applied independently to match the destination string. First, it checks that the first three digits are numbers, and then checks that the three digits are not "999 ". This pattern does not match a string with three digits before "foo" and followed by three characters not 999 and 6 Characters in total. For example, it does not match "123abcfoo ". The pattern that matches the string "123abcfoo" can be (? <= \ D {3 }...) (? <! 999) foo. In this case, check that the first three are digits, and then the second one is checked (the current match point) the first three characters are not "999"
Assertions can be nested with any complexity. For example (? <= (? <! Foo) bar) baz matches "bar" in front, but "bar" does not have "baz" in front of "foo ". Another mode (? <= \ D {3 }... (? <! 999) foo matches "foo" with three numbers followed by any character other than 999"
When the assertion sub-group is not captured, it cannot be modified by quantifiers, because it is meaningless to perform multiple assertions on the same thing. If all assertions contain one capture sub-group, they will be included for the purpose of capturing the Sub-Group count in the entire mode. However, substring capture can only be used for positive assertions, because it is meaningless for negative assertions.
The maximum number of sub-groups that can be included in assertion is 200.
One-time Sub-group
For repeated items with the maximum and minimum quantifiers restrictions at the same time, after the matching fails, another number of repetitions will be re-evaluated to determine whether the pattern can be matched. When the authors of the pattern clearly know that there is no problem with execution, it is useful to prevent such behavior by changing the matching behavior or making it fail to match earlier.
Consider an example. When the mode \ d + foo is applied to the 123456bar of the target row:
An error occurred while matching "foo" after six digits. During normal behavior, the matcher tries to make \ d + match only five digits and only four digits, try again before the final failure. A one-time Sub-Group provides a special meaning. When a part of the mode is matched, It is not re-evaluated, therefore, after the first failure to match "foo", the matching will immediately fail. The syntax symbol is another special bracket, with (?> Start, such as (?> \ D +) bar
This bracket provides a "Lock" for a part of the pattern. When it contains a match, it will prevent future pattern from returning to it after it fails. Backward Tracing is invalid here, and other work is carried out as usual
In other words, if the current matching point in the target string is an anchor, this type of substring matches a string equivalent to an independent pattern match.
A one-time sub-group is not a capture sub-group. In simple terms, the above example is to eat as many matching characters as possible. Therefore, even though \ d + and \ d +? The number of numbers to be matched is adjusted to match other parts of the pattern. (?> \ D +) but can only match the entire numerical sequence
This (syntax) structure can contain characters of any complexity or nesting
A one-time sub-group can be used in combination with the post-Zhan assertions to specify a valid match at the end of the target string. Consider when a simple pattern such as abcd $ is applied to an unmatched long string. Since the PCRE is processed from left to right during matching, the PCRE searches for each "a" from the target and then checks whether the remaining part of the matching mode is followed. If the mode is ^. * abcd $, then the initial. * The entire string will be matched first, but when it fails (because it is not followed by "a"), it will trace back all the matches and spit out the last character in sequence, the last 2nd characters. Find "a" in the entire string from the right to the left, so we cannot exit very well. However, if the mode of writing ^ (?>. *)(? <= Abcd). * This part is used only to match the entire string. Houzhan assertion tests the last four characters at the end of the string. If it fails, the matching immediately fails. For long strings, this mode will significantly improve the processing performance.
When a mode contains a sub-group that can be infinitely duplicated and has an infinite element in it, using a one-time sub-group is the only way to avoid some failure matching that consumes a lot of time. Mode (\ D + | <\ d +>) * [!?] Match a non-numeric character with no limit or followed by a <> closed numeric character! Or ?. When it matches, it runs quickly. However, if it is applied to "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa", it will consume a lot of time before reporting errors. This is because the string can be used for two repeated rules and must be allocated for both of them for an attempt. ([!?] Is used at the end of the example. Not a single character, because both PCRE and perl will optimize the Fast Error Reporting when the mode is last a single character. They record the individual characters that need to be matched at the end, and report an error quickly when they are not displayed in the string .) If the mode is modified to (?> \ D +) | <\ d +>) * [!?] An error is returned quickly.
Condition Sub-group
You can enable the matcher to match a sub-group in a condition based on the result of an asserted, or whether a previous capture sub-group matches or select between two optional sub-groups. The condition Sub-Group syntax is as follows:
(?(condition)yes-pattern)(?(condition)yes-pattern|no-pattern)
If the conditions are met, use yes-pattern. Otherwise, use no-pattern (if specified ). If there are more than two optional sub-groups, a compilation error will be generated.
There are two conditions. If the text is composed of numbers in the brackets of the condition, the condition is met when the child group represented by the number (Before) is matched (yes-pattern is used ). Consider the following pattern. To make it easy to read, add some blank characters (view the PCRE_EXTENDED option) and divide it into three parts :(\()? [^ ()] + (? (1 )\))
The first part of the pattern matches an optional left brace. If this character appears, set it as the capture substring of the first child group. The second part matches one or more non-parentheses. The third part is a condition Sub-Group, which will test whether the first sub-group matches. If yes, that is, the target string starts with braces and the condition is TRUE, if yes-pattern is used, an angle bracket must be matched. In other cases, since no-pattern does not appear, this sub-group does not match anything. In other words, this pattern matches a character sequence without parentheses or enclosed by parentheses.
If the conditional string (R) is used, it is used to obtain a recursive call to the mode or submode. The condition is always false at the "upper-level.
If the condition is not a numerical sequence or (R), it must be an asserted. The assertions here can be arbitrary, positive, negative, positive, and backward. Considering this mode, some white spaces are added to facilitate reading, and there are two optional paths in the second line.
(?(?=[^a-z]*[a-z])\d{2}-[a-z]{3}-\d{2} | \d{2}-\d{2}-\d{2} )
Conditional a positive asserted that matches an optional string of non-lowercase letters followed by a lowercase letter. In other words, it tests the target with at least one lower-case letter. If a lower-case letter is found, the target matches the first available branch. In other cases, the target matches the second branch. This pattern matches strings in two formats: dd-aaa-dd Or dd-dd. Aaa indicates lowercase letters, and dd indicates numbers.
Note
Character Sequence (? # Mark the start of a comment until a right parenthesis is encountered. Nested parentheses are not allowed. The characters in the comment are not used for matching as part of the pattern.
If the PCRE_EXTENDED option is set, the unescaped # characters outside the character class indicate that the remaining part of the line is a comment.
Recursive Mode
Infinite nested parentheses are allowed to match strings in parentheses. If recursion is not used, the best way is to use a pattern to match nested with a fixed depth. It cannot process any depth of nesting. Perl5.6 provides an experimental function that allows regular expression recursion. Special item (? R) provides this kind of recursive usage. This PCRE mode solves the parentheses problem (assuming that the PCRE_EXTENDED option is set, the blank characters are ignored): \ (?> [^ ()] +) | (? R ))*\).
First, it matches a left brace. Then it matches any number of non-parenthesis character sequences or a recursive match of the pattern itself (for example, a correct parenthesis substring), and finally matches a right parenthesis.
In this example, the pattern contains infinite duplicates, so a one-time sub-group is used to match non-Parentheses, which is very important when the pattern is applied to strings that do not match the pattern. For example, when it is applied to (aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa (), "mismatch" results will be generated quickly. However, if you do not use a one-time Sub-Group, this match will take a long time, because there are many ways to separate the target string by repeating the limit of + and, test all paths before reporting failure.
The final captured values of all capture sub-groups are captured from the recursion outermost submode. If the above pattern matches (AB (cd) ef), the final value of the capture sub-group is "ef", that is, the last value obtained by the top level. If additional parentheses are added, \ (?> [^ ()] +) | (? R) *) \), the captured string is the Matching content of the top-layer brackets "AB (cd) ef ". If there are more than 15 capturing parentheses in the mode, PCRE uses pcre_malloc to allocate additional memory during recursion to store data, and then releases them through pcre_free. If no memory can be allocated, it only saves the first 15 capturing parentheses, and the memory insufficiency error cannot be returned within recursion.
Starting from PHP4.3.3 ,(? 1 ),(? 2) can be used for Recursive subgroups. This can also be used to name sub-groups :(? P> name) or (? P & name ).
If the Recursive Sub-Group syntax is used outside the Sub-Group brackets it mentions (whether it is a sub-group number or sub-group name), this operation is equivalent to a sub-program in programming language. The previous example shows that the mode (sens | respons) eand \ 1ibility matches "sense and responsibility" and "response and responsibility", but does not match "sense and responsibility ". If the mode (sens | respons) e and (? 1) ibility substitution, which will match "sense and responsibility" just like matching the two strings ". This reference method is followed by matching the referenced submodel.
The maximum length of the target string is the maximum positive integer that can be stored in int variables. However, PCRE uses recursion to process subgroups and infinite duplicates. This means that the stack space available in some modes may be limited by the target string.
Performance
Some items in the schema may be more efficient than others. For example, the use of character classes such as [aeiou] is more efficient than the optional path (a | e | I | o | u. Generally, it is the most efficient to describe the requirement with the simplest possible structure.
When a mode starts with. * And the PCRE_DOTALL option is set, the mode is implicitly anchored through PCRE because it can match the start of a string. However, if PCRE_DOTALL is not set, PCRE cannot perform this optimization because. metacharacters cannot match line breaks. If the target string contains line breaks, the pattern may start matching from the end of a line break, rather than the start position. For example, the mode (. *) second matches the target string "first \ nandsecond" (\ n is a line break). The first capture Sub-Group result is "and ". To do this, PCRE tries to match each line break in the target string
If you use the pattern to match the target string without a linefeed, you can specify the position by setting PCRE_DOTALL or starting with the pattern ^. * to get the best performance. This saves the PCRE time to start scanning and searching for linefeeds along the target string.
Infinite repeated nesting in careful mode. This may cause a long running time when applying a unmatched string. Consider the mode fragment (a + )*
In some simple cases, optimization is like (a +) * B followed by using the original string .. Before getting started with the formal match, PCRE checks whether the target string contains the "B" character. If not, it immediately fails. However, this optimization is not available when there are no original characters. You can compare the behavior differences between (a +) * \ d and the above pattern. The former reports failures almost immediately when applying a string consisting of "a" to the entire line, while the latter reports time consumption when the target string is longer than 20 characters.