PHP Regular matching Chinese UTF8 encoding/^[\x{4e00}-\x{9fa5}]+$/____ encoding
Source: Internet
Author: User
In JavaScript, it's easy to tell if a string is Chinese. Like what:
var str = "PHP programming";
if (/^[\u4e00-\u9fa5]+$/.test (str)) {
Alert ("The string is all in Chinese");
} else {
Alert ("This string is not all Chinese");
}
Taken for granted, in PHP to determine whether the string is Chinese, it will follow this idea:
<?php
$STR = "PHP programming";
if (Preg_match ("/^[\u4e00-\u9fa5]+$/", $str)) {
Print ("The string is all Chinese");
} else {
Print ("This string is not all Chinese");
}
?>
However, it will soon be found that PHP does not support this expression, the error:
Warning:preg_match () [Function.preg-match]: compilation Failed:pcre does not support \l, \l, \ n, \u, or \u at offset 3 I n test.php on line 3
Just started looking up a lot of times from Google, and wanted to start with the PHP regular expression for hexadecimal data
The expression of a breakthrough, found in PHP, is used to represent hexadecimal data \x. So
Transform into the following code:
$STR = "PHP programming";
if (Preg_match ("/^[\x4e00-\x9fa5]+$/", $str)) {
Print ("The string is all Chinese");
} else {
Print ("This string is not all Chinese");
}
Seemingly no error, the results of the judgment is correct, but the $STR replaced by "programming" two words, the result is
Or "The string is not all Chinese", it seems that the judgment is not accurate enough.
Later ran back to Baidu search "PHP matching Chinese characters UTF 8", found that the article is more than the degree of match Google's higher,
It seems that Baidu's "more understand Chinese" is still to a certain extent correct. In the second article "★★★ Seek UTF8
To match the Chinese characters of the regular, online and so on ... "see the following elements:
Landlord Zhiin (┈jcan┈) 2006-11-15 15:59:30 in WEB development/PHP Questions
To find the UTF8 matching Chinese characters, excluding full-width characters and special symbols!
Only regular matching full-width characters can be found on the net: ^[\x80-\xff]*^/
[\u4e00-\u9fa5] can match Chinese, but PHP does not support
Depressed in ....
1/F pleasedotellmewhy (Allah bless you!) reply to 2006-11-15 16:04:55 score 11
Chr (0XA1). '-' . Chr (0xff) can match all Chinese, but don't know what to do under UTF-8! Top
2/F Zhiin (┈jcan┈) reply to 2006-11-15 16:11:34 score 0
Even under GB2312, Chr (0XA1). '-' . Chr (0xff) is not right
It also matches the full-width symbols in the top.
3/F xuzuning (NAG) back to 2006-11-15 16:19:56 score 90
Pattern modifier: U
After trying each of these clues, it turns out that, as they say, it might have something to do with coding,
So you need to know about the pattern modifier-so keep searching for Baidu.
In a "pattern modifier" article, read:
U (PCRE_UTF8)
This modifier enables an additional feature that is incompatible with Perl in a PCRE. The pattern string is treated as UTF-8.
This modifier is available under Unix from PHP 4.1.0 and is available under Win32 from PHP 4.2.3.
Example:
Preg_match ('/[\x{2460}-\x{2468}]/u ', $str); Matching inner code Chinese characters
In the way he provided, the code was as follows:
$STR = "PHP programming";
if (Preg_match ("/^[\x{2460}-\x{2468}]+$/u", $str)) {
Print ("The string is all Chinese");
} else {
Print ("This string is not all Chinese");
}
Find out whether or not to judge the Chinese is still abnormal. However, since the hexadecimal data represented by \x,
Why and JS inside the scope of the \X4E00-\X9FA5 is not the same. So I switched to the bottom code:
$STR = "PHP programming";
if (Preg_match ("/^[\x4e00-\x9fa5]+$/u", $str)) {
Print ("The string is all Chinese");
} else {
Print ("This string is not all Chinese");
}
Originally thought the thing that definitely succeeds, unexpectedly, warning again produce:
Warning:preg_match () [Function.preg-match]: compilation Failed:invalid UTF-8 string at offset 6 into test.php on line 3
There seems to be a wrong way of saying it, and then I compare the way the article was expressed,
To "4e00" and "9fa5" each side with "{" and "}" wrapped up, ran again, found really accurate:
$STR = "PHP programming";
if (Preg_match ("/^[\x{4e00}-\x{9fa5}]+$/u", $str)) {
Print ("The string is all Chinese");
} else {
Print ("This string is not all Chinese");
}
Know the final correct expression--/^[\x{4e00}-\x{9fa5}]+$/u of the Chinese characters with regular expressions in PHP utf-8 encoding,
So I used this expression to Baidu search, found that there is really someone else came up with such a correct conclusion, but through
The conventional way is hard to find, and only one--"delete Chinese characters with positive", seems to be on the internet for
The selection of the correctness of information should be strengthened urgently.
PS: Google will not give up, but also search for a while, and found an article "PHP Common Class",
Or in the Baidu space, hehe, interesting.
--------------------------------------------------------------------------------------------------------------- -------------------
Reference to the above article wrote the following section of the test code (copy the following code saved into a. php file)
<?php
$action = Trim ($_get[' action '));
if ($action = = "Sub")
{
$str = $_post[' dir '];
if (!preg_match ("/^[". Chr (0XA1). " -". Chr (0xff)." a-za-z0-9_]+$/", $str))//gb2312 Chinese character alphanumeric underline regular expression
if (!preg_match ("/^[\x{4e00}-\x{9fa5}a-za-z0-9_]+$/u", $str))//utf-8 Chinese character alphanumeric underline regular expression
{
echo "<font color=red> you entered [". $str. "] Contain illegal characters </font> ";
}
Else
{
echo "<font color=green> you entered [". $str. "] Perfectly lawful, through the!</font> ";
}
}
?>
<form method= "POST" action= "Action=sub" >
Input characters (numbers, letters, Chinese characters, underscores):
<input type= "text" name= "dir" value= "" >
<input type= "Submit" value= "submitted" >
</form>
The content source of this page is from Internet, which doesn't represent Alibaba Cloud's opinion;
products and services mentioned on that page don't have any relationship with Alibaba Cloud. If the
content of the page makes you feel confusing, please write us an email, we will handle the problem
within 5 days after receiving your email.
If you find any instances of plagiarism from the community, please send an email to:
info-contact@alibabacloud.com
and provide relevant evidence. A staff member will contact you within 5 working days.
A Free Trial That Lets You Build Big!
Start building with 50+ products and up to 12 months usage for Elastic Compute Service