One: If it is a small file, can be read into the array at once, using a convenient array count function for Word frequency statistics (assuming that the contents of the file are space-separated words):
<?php $str = file_get_contents ("/path/to/file.txt");//get string from File preg_match_all ("/\b (\w+[-]\w+ )| (\w+) \b/", $str, $r); Place words into array $r-this includes hyphenated words $words = array_count_values (Array_map ("Strtolower", $r [0] )); Create new Array-with case-insensitive Count arsort ($words);//order from high to low Print_r ($words)
Second: If it is a large file, read into the memory is not appropriate, you can use the following methods:
<?php $filename = "/path/to/file.txt"; $handle = fopen ($filename, "R"); if ($handle = = = False) { exit; } $word = ""; while (false!== ($letter = fgetc ($handle))) { if ($letter = = ") { $results [$word]++; $word = ""; } else { $word. = $letter; } } Fclose ($handle); Print_r ($results);
Linux Commands Classic interview: The top 10 most frequently occurring words in a statistical file
Using the Linux command or the shell implementation: The file words store English words in the form of one English word per line (words can be duplicated), counting the top 10 occurrences of the file.
Cat Words.txt | Sort | uniq-c | SORT-K1,1NR | Head-10
The main investigation of the use of the sort, uniq command, the relevant explanation is as follows, the command and parameters of the detailed instructions please self through the man view, a brief introduction of the above instructions of the functions of each part:
Sort : sorting words
uniq-c: displays unique rows and the number of occurrences of the bank in the file at the beginning of each line
SORT-K1,1NR: sorted by first field, numerically, and in reverse order
head-10: take the first 10 rows of data
PHP: Calculate how often words appear in a file or array