雜湊表對字串的高效處理

來源:互聯網
上載者:User

雜湊表對字串的高效處理

        雜湊表(散列表)是一種非常高效的尋找資料結構,在原理上也與其他的尋找不盡相同,它迴避了關鍵字之間反覆比較的繁瑣,而是直接一步到位尋找結果。當然,這也帶來了記錄之間沒有任何關聯的弊端。應該說,散列表對於那些尋找效能要求高,記錄之間關係無要求的資料有非常好的適用性。注意對散列函數的選擇處理衝突的方法

        Hash表是使用 O(1)
時間進行資料的插入、刪除和尋找,但是 hash 表不保證表中資料的有序性,這樣在 hash 表中尋找最大資料或者最小資料的時間是 O(N) 。

 

/* 字串中完成過濾重複字元的功能,

【輸入】:1.常字串;2.字串長度;3.【out】用於輸出過濾後的字串.

【輸出】:過濾後的字串。

*/

思路1, 迴圈判定法。第1步,先記錄字串中第1個字元;第2步,然後從第2個字元開始,判定其和其前面的字元是否相同,不相同的話,則統計進去;相同的話則繼續遍曆,直到字串末尾(遇到’\0’)。時間複雜度:O(n2)。

思路2, 雜湊表過濾法。第1步,初始化一個雜湊表,用以儲存字元(key)及字元出現的次數;第2步,遍曆雜湊表,進行統計計數;第3步,輸出統計次數為1及統計次數多餘1的(輸出1次)。時間複雜度:O(n)。

//迴圈判定法過濾掉重複字元

void stringFilter(const char*pInputStr, long lInputLen, char *pOutputStr){       if(pInputStr== NULL || lInputLen == 0 || pOutputStr == NULL)       {              return;       }             intnCnt = 0;       *pOutputStr= pInputStr[0];            //先處理第一個       ++nCnt;             intnNotEqualCnt = 0;                 //統計計數       for(inti = 1; i < lInputLen; i++)       {              nNotEqualCnt= 0;              for(intj = i-1; j >=0; j--)              {                     if(pInputStr[i]!= pInputStr[j])                     {                            ++nNotEqualCnt;                     }              }                           if(nNotEqualCnt== i)  //和前面的都不一樣.              {                     pOutputStr[nCnt++]= pInputStr[i];              }                    }//endfor       pOutputStr[nCnt]= '\0';}

//雜湊表法過濾字串中的重複字元

void stringFilterFast(const char*pInputStr, long lInputLen, char *pOutputStr){       charrstChar = '\0';       boolbNotRepeatFound = false;       constunsigned int size = 256;       unsignedint hashTable[size];       constchar* pHashKey = pInputStr;       intoutPutCnt = 0;             if(pInputStr== NULL)       {              return;       }             //初始化雜湊表       for(unsignedint i = 0; i < size; i++)       {              hashTable[i]= 0;       }             //將pString讀入到雜湊表中       while(*pHashKey!= '\0')       {              cout<< *pHashKey << "\t";              hashTable[*pHashKey]++;    //統計計數              pHashKey++;       }            //讀取雜湊表,對只出現1次的進行儲存,對出現多次的進行1次儲存。       pHashKey= pInputStr;       while(*pHashKey!= '\0')       {              if((hashTable[*(pHashKey)])== 1)   //僅有一次,              {                     pOutputStr[outPutCnt++]= *pHashKey;              }              elseif((hashTable[*(pHashKey)]) > 1) // 多餘一次,統計第一次              {                     pOutputStr[outPutCnt++]= *pHashKey;                     hashTable[*(pHashKey)]= 0;              }              pHashKey++;       }       pOutputStr[outPutCnt]= '\0'; } int main(){       constchar* strSrc = "desdefedeffdsswwwwwwwwwwdd";//"desdefedeffdssw";       char*strRst =new char[strlen(strSrc)+1];       stringFilter(strSrc,strlen(strSrc), strRst);       cout<< strRst << endl;       return0;}

//雜湊表法尋找字串中第一個不重複的字元

【功能】:尋找字串中第一個不重複的字元。

【輸入】:字串。

【輸出】:第一個不重複的字元。

時間複雜度O(n),思路類似於上面的雜湊表過濾法

char FirstNotRepeatingChar(constchar* pString){       charrstChar = '\0';       boolbNotRepeatFound = false;       constunsigned int size = 256;       unsignedchar hashTable[size];       constchar* pHashKey = pString;        if(pString== NULL)       {              returnrstChar;       }        //初始化雜湊表       for(unsignedint i = 0; i < size; i++)       {              hashTable[i] = 0;       }             //將pString存入到雜湊表中       while(*pHashKey!= '\0')       {              hashTable[*(pHashKey++)]++;    //統計計數       }        //讀取雜湊表,找到第一個=1的字元,bNotRepeatFound用於尋找。.       pHashKey= pString;       while(*pHashKey!= '\0')       {              if((hashTable[*(pHashKey)]) == 1)              {                     bNotRepeatFound= true;                     rstChar= *pHashKey;                     break;              }              pHashKey++;       }        if(bNotRepeatFound)       {              cout<< "The first not Repeate char is " << rstChar <<endl;       }       else       {              cout<< "The first not Repeate char is not Exist " << endl;       }        returnrstChar;} int main(){       constchar* strSrc = "google";       constchar* strSrc2 = "yyy@163.com";       constchar* strSrc3 = "aabbccddeeff";       constchar* strsrc4 = "11111111";                                                                                                                            constchar* strArray[4] = {strSrc, strSrc2, strSrc3, strsrc4};       for(inti = 0; i < 4; i++)      {              FirstNotRepeatingChar(strArray[i]);       }        return0;}

舉一反三:【百度面試題】對於一個海量的檔案中儲存著不同的URL,用最小的時間複雜度去除重複的URL。可借鑒字串處理的雜湊表過濾法。不過,這裡的大檔案等價於之前的字串,這裡的URL等價於之前的字元。

聯繫我們

該頁面正文內容均來源於網絡整理,並不代表阿里雲官方的觀點,該頁面所提到的產品和服務也與阿里云無關,如果該頁面內容對您造成了困擾,歡迎寫郵件給我們,收到郵件我們將在5個工作日內處理。

如果您發現本社區中有涉嫌抄襲的內容,歡迎發送郵件至: info-contact@alibabacloud.com 進行舉報並提供相關證據,工作人員會在 5 個工作天內聯絡您,一經查實,本站將立刻刪除涉嫌侵權內容。

A Free Trial That Lets You Build Big!

Start building with 50+ products and up to 12 months usage for Elastic Compute Service

  • Sales Support

    1 on 1 presale consultation

  • After-Sales Support

    24/7 Technical Support 6 Free Tickets per Quarter Faster Response

  • Alibaba Cloud offers highly flexible support services tailored to meet your exact needs.