Jieba word segmentation,
Jieba is written in python for word splitting, with clear code and good scalability. You can easily write your own code for magic modification with the idea of improving jieba.
Basic Idea of jieba Word Segmentation
The jieba word segmentation algorithm is used to process both indexed and unindexed words. The processing method is simple. Of course, an overly simple algorithm is also one of the reasons that restricts the recall rate.
The main solution is as follows:
1.w.dict.txt
2. Build the DAG of the sentence from the memory Dictionary (directed acyclic graph)
3. For words not included in the dictionary, use the viterbi Algorithm of HMM to try word segmentation.
4. Use dp to find the maximum probability path of the DAG after all the words that have been included and those that have not been included are segmented.
5. Output word segmentation results
Jieba Word Segmentation: to quickly index a dictionary to accelerate word segmentation, A dict is constructed using a prefix array to store the dictionary.
In earlier versions of jieba, jieba uses the trie tree data structure for storage. In fact, it is very redundant to use the trie tree in python.
Trie tree
The trie tree is also called a dictionary tree. It is a common data structure used for fast string matching in a string list. The core idea is to normalize words with a public prefix to a tree to reduce the query time complexity. The main drawback is that the memory usage is too large.
The trie tree is constructed as follows:
A trie tree constructed with and as cn com is as follows:
Traverse each row of files. For each letter of each word, check whether the trie tree (trie and p variables) exists. If so, it is mounted to the following. If not, create a new subtree.
Jieba uses python dict to store trees. This is also a common practice of python for tree data structures.
Trie tree Problems
Originally, jieba could use the trie tree as its starting point. It could use space for time to speed up word segmentation and accelerate full splitting. However, the problem is that python's dict uses a hash table for native implementation. It takes almost O (1) Time to obtain words in dict. Therefore, using the trie tree is actually a way to avoid duplication.
Prefix Array
In a PR (https://github.com/fxsjy/jieba/pull/187) in 2014, the submitter changed the trie tree to a prefix array, greatly reducing memory usage and speeding up search.
Now, the operations of jieba word segmentation for dictionaries are changed to a word: freq structure, which is stored in lfreq. The specific operations are as follows:
1. For each indexed word, if it is in lfreq, the word frequency is accumulated. If not, lfreq is added.
2. Perform the previous operation on all prefixes of the indexed word, such as the word 'cat', and perform the first operation on c, ca, and cat respectively. All prefixes except the word itself are initially 0.
def gen_pfdict(self, f): lfreq = {} ltotal = 0 f_name = resolve_filename(f) for lineno, line in enumerate(f, 1): try: line = line.strip().decode('utf-8') word, freq = line.split(' ')[:2] freq = int(freq) lfreq[word] = freq ltotal += freq for ch in xrange(len(word)): wfrag = word[:ch + 1] if wfrag not in lfreq: lfreq[wfrag] = 0 except ValueError: raise ValueError( 'invalid dictionary entry in %s at Line %s: %s' % (f_name, lineno, line)) f.close() return lfreq, ltotal
It is a simple practice. However, it makes full use of the python dict type and improves the efficiency a lot.
Jieba word segmentation can be selected in multiple modes. Optional modes include:
Full splitting Mode
Exact Mode
Search engine Mode
# Encoding = utf-8import jiebaseg_list = jieba. cut ("I came to Beijing Tsinghua University", cut_all = True) print ("Full Mode:" + "/". join (seg_list) # Full mode seg_list = jieba. cut ("I came to Beijing Tsinghua University", cut_all = False) print ("Default Mode:" + "/". join (seg_list) # exact mode seg_list = jieba. cut ("he has come to Netease hang Yan building") # The default is the exact mode print (",". join (seg_list) seg_list = jieba. cut_for_search ("James graduated from the Institute of Computing Science of the Chinese Emy of sciences and later studied at Kyoto University in Japan") # search engine mode print (",". join (seg_list ))
Output result:
[Full mode]: I/come/Beijing/Tsinghua University/Huada/University [exact mode]: I/come/Beijing/Tsinghua University [New Word Recognition]: He, come, Netease, hangyan, building (here, "hangyan" is not in the dictionary, but it is also identified by Viterbi algorithms) [Search Engine mode]: James, master's degree, graduated from, China, science, college, Emy of sciences, Chinese Emy of Sciences, Institute of computing, and later in, Japan, Kyoto University, Kyoto University, Japan for further studies