Jieba word segmentation,

Source: Internet
Author: User

Jieba word segmentation,

Jieba is written in python for word splitting, with clear code and good scalability. You can easily write your own code for magic modification with the idea of improving jieba.

Basic Idea of jieba Word Segmentation

The jieba word segmentation algorithm is used to process both indexed and unindexed words. The processing method is simple. Of course, an overly simple algorithm is also one of the reasons that restricts the recall rate.

The main solution is as follows:
1.w.dict.txt
2. Build the DAG of the sentence from the memory Dictionary (directed acyclic graph)
3. For words not included in the dictionary, use the viterbi Algorithm of HMM to try word segmentation.
4. Use dp to find the maximum probability path of the DAG after all the words that have been included and those that have not been included are segmented.
5. Output word segmentation results

Jieba Word Segmentation: to quickly index a dictionary to accelerate word segmentation, A dict is constructed using a prefix array to store the dictionary.
In earlier versions of jieba, jieba uses the trie tree data structure for storage. In fact, it is very redundant to use the trie tree in python.

Trie tree
The trie tree is also called a dictionary tree. It is a common data structure used for fast string matching in a string list. The core idea is to normalize words with a public prefix to a tree to reduce the query time complexity. The main drawback is that the memory usage is too large.
The trie tree is constructed as follows:
A trie tree constructed with and as cn com is as follows:

Traverse each row of files. For each letter of each word, check whether the trie tree (trie and p variables) exists. If so, it is mounted to the following. If not, create a new subtree.

Jieba uses python dict to store trees. This is also a common practice of python for tree data structures.

Trie tree Problems

Originally, jieba could use the trie tree as its starting point. It could use space for time to speed up word segmentation and accelerate full splitting. However, the problem is that python's dict uses a hash table for native implementation. It takes almost O (1) Time to obtain words in dict. Therefore, using the trie tree is actually a way to avoid duplication.

Prefix Array

In a PR (https://github.com/fxsjy/jieba/pull/187) in 2014, the submitter changed the trie tree to a prefix array, greatly reducing memory usage and speeding up search.

Now, the operations of jieba word segmentation for dictionaries are changed to a word: freq structure, which is stored in lfreq. The specific operations are as follows:
1. For each indexed word, if it is in lfreq, the word frequency is accumulated. If not, lfreq is added.
2. Perform the previous operation on all prefixes of the indexed word, such as the word 'cat', and perform the first operation on c, ca, and cat respectively. All prefixes except the word itself are initially 0.

def gen_pfdict(self, f):        lfreq = {}        ltotal = 0        f_name = resolve_filename(f)        for lineno, line in enumerate(f, 1):            try:                line = line.strip().decode('utf-8')                word, freq = line.split(' ')[:2]                freq = int(freq)                lfreq[word] = freq                ltotal += freq                for ch in xrange(len(word)):                    wfrag = word[:ch + 1]                    if wfrag not in lfreq:                        lfreq[wfrag] = 0            except ValueError:                raise ValueError(                    'invalid dictionary entry in %s at Line %s: %s' % (f_name, lineno, line))        f.close()        return lfreq, ltotal

It is a simple practice. However, it makes full use of the python dict type and improves the efficiency a lot.

Jieba word segmentation can be selected in multiple modes. Optional modes include:

Full splitting Mode
Exact Mode
Search engine Mode

# Encoding = utf-8import jiebaseg_list = jieba. cut ("I came to Beijing Tsinghua University", cut_all = True) print ("Full Mode:" + "/". join (seg_list) # Full mode seg_list = jieba. cut ("I came to Beijing Tsinghua University", cut_all = False) print ("Default Mode:" + "/". join (seg_list) # exact mode seg_list = jieba. cut ("he has come to Netease hang Yan building") # The default is the exact mode print (",". join (seg_list) seg_list = jieba. cut_for_search ("James graduated from the Institute of Computing Science of the Chinese Emy of sciences and later studied at Kyoto University in Japan") # search engine mode print (",". join (seg_list ))

Output result:

[Full mode]: I/come/Beijing/Tsinghua University/Huada/University [exact mode]: I/come/Beijing/Tsinghua University [New Word Recognition]: He, come, Netease, hangyan, building (here, "hangyan" is not in the dictionary, but it is also identified by Viterbi algorithms) [Search Engine mode]: James, master's degree, graduated from, China, science, college, Emy of sciences, Chinese Emy of Sciences, Institute of computing, and later in, Japan, Kyoto University, Kyoto University, Japan for further studies

  

Contact Us

The content source of this page is from Internet, which doesn't represent Alibaba Cloud's opinion; products and services mentioned on that page don't have any relationship with Alibaba Cloud. If the content of the page makes you feel confusing, please write us an email, we will handle the problem within 5 days after receiving your email.

If you find any instances of plagiarism from the community, please send an email to: info-contact@alibabacloud.com and provide relevant evidence. A staff member will contact you within 5 working days.

A Free Trial That Lets You Build Big!

Start building with 50+ products and up to 12 months usage for Elastic Compute Service

  • Sales Support

    1 on 1 presale consultation

  • After-Sales Support

    24/7 Technical Support 6 Free Tickets per Quarter Faster Response

  • Alibaba Cloud offers highly flexible support services tailored to meet your exact needs.