標籤:
題目重現
In 1953, David A. Huffman published his paper “A Method for the Construction of Minimum-Redundancy Codes”, and hence printed his name in the history of computer science. As a professor who gives the final exam problem on Huffman codes, I am encountering a big problem: the Huffman codes are NOT unique. For example, given a string “aaaxuaxz”, we can observe that the frequencies of the characters ‘a’, ‘x’, ‘u’ and ‘z’ are 4, 2, 1 and 1, respectively. We may either encode the symbols as {‘a’=0, ‘x’=10, ‘u’=110, ‘z’=111}, or in another way as {‘a’=1, ‘x’=01, ‘u’=001, ‘z’=000}, both compress the string into 14 bits. Another set of code can be given as {‘a’=0, ‘x’=11, ‘u’=100, ‘z’=101}, but {‘a’=0, ‘x’=01, ‘u’=011, ‘z’=001} is NOT correct since “aaaxuaxz” and “aazuaxax” can both be decoded from the code 00001011001001. The students are submitting all kinds of codes, and I need a computer program to help me determine which ones are correct and which ones are not.
Input Specification
Each input file contains one test case. For each case, the first line gives an integer N (2≤N≤63), then followed by a line that contains all the N distinct characters and their frequencies in the following format:
c[1] f[1] c[2] f[2] ... c[N] f[N]
where c[i] is a character chosen from {‘0’ - ‘9’, ‘a’ - ‘z’, ‘A’ - ‘Z’, ‘_’}, and f[i] is the frequency of c[i] and is an integer no more than 1000. The next line gives a positive integer M (≤1000), then followed by M student submissions. Each student submission consists of N lines, each in the format:
c[i] code[i]
where c[i] is the i-th character and code[i] is an non-empty string of no more than 63 ‘0’s and ‘1’s.
Output Specification
For each test case, print in each line either “Yes” if the student’s submission is correct, or “No” if not.
Note: The optimal solution is not necessarily generated by Huffman algorithm. Any prefix code with code length being optimal is considered correct.
Sample Input
7A 1 B 1 C 1 D 3 E 3 F 6 G 64A 00000B 00001C 0001D 001E 01F 10G 11A 01010B 01011C 0100D 011E 10F 11G 00A 000B 001C 010D 011E 100F 101G 110A 00000B 00001C 0001D 001E 00F 10G 11
Sample Output
YesYesNoNo
題目大意
給定詞頻序列,判定給定的若干組編碼方式是否與Huffman編碼等效。
要點有兩個,一個是要求編碼不產生歧義,另一個是總編碼長度最短。
解法求詞頻序列的Huffman編碼長度
單獨將詞頻序列提出來,可以產生一個唯一的Huffman編碼長度,這也是最優的長度。
按照Huffman演算法,每次提取兩個最小的,合并,最後就可以得到這個最短長度。
這裡可以使用插入排序,也可以直接用優先隊列加速。
判定是否有編碼歧義
根據給出的編碼方式,構造一個Trie Tree(字典樹)。這個字典樹的每個節點要存:
bool isVisited; // 是否被訪問過bool isMarked; // 是否被標記佔用Trie *next[2]; // 指向下一級節點
當按照字串構造時,注意沿途做如下標記:
每當訪問到一個節點,isVisited = true;
每當抵達終點,使isMarked = true;
即:
- 具有
isVisited 標記的節點不能是新編碼的終點,否則新編碼就是某個編碼的首碼子碼。
- 當經過
isMarked 標記時,中斷,否則某個編碼一定是新編碼的首碼子碼。
這樣就可以保證所有的終點都是葉子節點了。
代碼實現
PTA Huffman Codes With Trie Tree
PTA Huffman Codes