The standardization research of the series of search engines for information retrieval trends is carried by: China indexing Institute)
Source: Internet
Author: User
// The text is relatively simple, but it can be summarized by yourself. The following documents can also be referenced.China Index Institute http://www.cnindex.fudan.edu.cnHttp://www.cnindex.fudan.edu.cn/zgsy/2006n3/heshaohua.htmResearch on standardization of search enginesHe Shaohua Sun Chen (School of information management, Wuhan University, Wuhan 430072, China)
AbstractYesThis article introduces the concept of a search engine, analyzes the current situation of the search engine, focuses on the existing problems of the current search engine, points out the shortcomings, the aim is to call for the establishment of appropriate standards for search engines based on actual conditions as soon as possible to effectively solve the problem of co-construction and sharing of network information.
KeywordsSearch engine Standardization
1IntroductionAfter years of development, search engines become more and more powerful and provide more comprehensive services. Their goal is to become the preferred Internet portal site for users, not only does it provide simple query functions. Integration, industrialization, intelligence, and multi-language will be the development direction of the new generation of search engines. In the future, the search engine will be more personalized and intelligent, with more information, faster search speed, higher accuracy, and more satisfying user needs. However, to achieve this goal, we must make great efforts to establish a standardized system for search engines.
2Search engine conceptNetwork search engine, English name: search engine, generally, it refers to a special site or server that provides online information resource retrieval and navigation services to network users through the hypermedia technology and the Internet [1]. A search engine task consists of two processes: one is on the server side, that is, the service provider collects online information, such as web pages and URLs, and non-www bbs, FTP, news, etc., search and analysis indexing, the establishment of the corresponding index database, and automatic tracking of information source changes, constantly update index records, the indexing database is regularly maintained. Second, when the user asks for retrieval, the server searches for its own information index database and sends it to the user. The former can be called the information indexing process, and the latter can be called the retrieval service process.
3Analysis of the current situation of search engines3.1 The classification of search engines currently has a large number of sites providing search services on the Internet. Their search engines vary in terms of the scope and method of indexing, each of which has its own merits. Based on its technical principles, it can be divided into three categories: robot-based search engines, directory-based search engines, and meta-search engines. (1) robot-based search engines. The search engine's robot (spider or crawler) is a set of known documents that roam online along the world wide web hyperlink. It uses the breadth-first and depth-first algorithms for search. Once a new website is found, the robot records the URL for further access, so that the new website is indexed until it is not found, relevant information is extracted from the title, meta tags, and other information that represents special documents as index items and added to the index database for user query. (2) directory-based search engines. The directory-based search engine classifies the collected information into a class [2]. A typical directory-based search engine is Yahoo. (3) Meta Search Engine. When a user queries, the Meta Search Engine calls multiple other independent search engines and processes the results obtained from multiple search engines, such as deleting duplicate results, test links, and sorting results. Such search engines do not need to create databases for websites, webpages, FTP, and other resources on the Internet for organization and maintenance, you only need to store the information of the connected site. The meta-search engine is easy to design, but the network load is too large. Typical meta-search engines include metacrawler. 3.2 The Status Quo of search engines at home and abroad in 1993 the British nexor company developed the first batch of online search tools aliweb. Today, foreign search engines have developed rapidly and are relatively mature, in addition, it has mature technologies in terms of database indexing scope, online information organization, search performance, interface friendliness, and result feedback. Compared with foreign search engines, Chinese search engines started late in the second half of 1997, currently, the designs of Baidu, Youyou, Yahoo, huahaogunjing, compass, and jubaobao have been used by several large websites in China. In terms of the number of indexed items, foreign search engines can search for hundreds of millions of webpages, but currently there are only more than 2000 in China, and there is a lack of full-network search engines, and different search engines have a high repetition rate, the distribution of professional network information is uneven, and the speed is also different. The Retrieval Technology and retrieval results are not satisfactory [3]. In addition to database capacity differences, Chinese search engines cannot compare their relevance sorting functions with foreign search engines. For relevance sorting of search results, Google creates a page-level technology based on citation analysis, page rank, by analyzing the number of links to the search results and the quality of the source links, the relevance sorting of the search results is achieved, which plays a greater role in the network metrology research, especially the network site sorting. Famous foreign search engines rely mainly on search engines to provide diversified services to attract more users, so as to get more advertising revenue. However, domestic search engines are not doing enough, many search engines focus on advertising rather than services, leading to poor service quality. Currently, the search engines used by major portals in China are generally written in Chinese directly or indirectly using English search software. These software did not take into account China's usage habits during the design and cannot be fully localized. Therefore, there are many limitations. To sum up, although Chinese search engines are continuously developing, they are still far behind the advanced level in foreign countries, and there are still many shortcomings and shortcomings.
4Search engine problemsSearch engines are currently the most important network information retrieval tools, its technology involves information retrieval, artificial intelligence, computer networks, distributed processing, databases, data mining, digital libraries, natural language processing, and other fields. Therefore, to explore the shortcomings of search engines, we can start from the relevant aspects. 4.1 different user query interfaces are used to enter user queries, display query results, and provide user relevance feedback mechanisms. They are mainly used to facilitate users' use of search engines, efficient and timely retrieval of information from search engines in multiple ways [4]. For user query interfaces, various search engines provide different implementation methods, either technically or in a method. Its ease of use and user friendliness need to be further improved. At present, some companies and organizations are considering developing standards for query options. 4.2 The indexing depth of information by the search engine is insufficient. Currently, network information mining is based on the form, such as keywords and titles. The obtained information and set requirements are only simple matching. Search engine search results often only provide some linear URLs and web page information including keywords, especially the retrieval of specific literature databases seems powerless [5]. A computer cannot understand the text. It must represent the content of a web page in binary format. Currently, most robot search engines on the Internet are used to filter, describe, and index pages based on the frequency and location of words and phrases appearing on pages, thus forming an index database for users to query. However, images on pages are not indexed. In addition, dynamic Web pages are not indexed due to their dynamic nature and structure. Therefore, for Chinese search engines, we need to use network data mining and knowledge extraction to analyze information content and their relationships, increase the indexing depth, and increase the multimedia retrieval function, that is, the discrete media represented by text information and the content of continuous media represented by images and sounds should be retrieved. Due to the wide coverage of multimedia information and complex objects, a high-level retrieval mechanism must be established to retrieve multimedia and its components in a unified manner. 4.3 The Retrieval languages of search engines are not flexible, accurate, and standardized. Artificial and controlled languages are currently the mainstream languages for information retrieval, however, users are required to use standard words to accurately express the retrieval content they are not familiar with. This increases the Learning Burden of the searcher and reduces the retrieval efficiency. Natural Language can better express Users' Query requirements, improve query accuracy, and facilitate the interaction between search engines and users. Therefore, natural language-based retrieval is an inevitable trend of development. in foreign countries, the introduction of natural language processing into Information Retrieval has been applied in theoretical research, and it is still in the theoretical research phase in China. However, natural language retrieval lacks standards. Therefore, how to establish a unified standard and provide high-quality retrieval through natural language interfaces is a problem that is constantly being explored in the Field of Information Review. 4.4 restrictions on a single search engine because the amount of information on the Internet is growing, it is impossible for a single search engine to include information resources of the entire network. Therefore, you must use all search engines to find the desired information. In addition, each engine often overwrites each other, and the user will repeatedly find the same information. Therefore, if an all-in-one search service is provided, Internet users only need to enter a query target once during the search process to obtain various associated query results on the same interface, which will be the future development trend, meta Search Engine and distributed search engine can solve this problem well. The unification of 4.5 languages is incompatible with the deepening of international exchanges. It is impossible to use only one English language as the search engine language. We should gradually develop multilingual search engines. For example, Google's server will automatically identify the country to which the computer belongs, and display the text in the country for non-English users to use [6]. In addition, because the cultural traditions, ways of thinking, and living habits of different countries in the world are different, search engines must be compatible, this requires developers to grasp the balance between standards and compatibility. 4.6 The absence of a unified web version of the word search table and Word Segmentation Word Table can make the search engine more intelligent. It is especially useful when using natural languages. In addition, Chinese words do not contain spaces and need to be segmented manually. In addition, Chinese words have a lot of ambiguity, there can be many results. Therefore, the absence of a unified word search table and Word Segmentation vocabulary will cause inconvenience to the server and the user. Therefore, we must create a unified standard search and Word Segmentation vocabulary for the search engine as soon as possible. 4.7 lack of evaluation criteria for result feedback currently, most search engines only provide links and brief descriptions of sites in result feedback, and lack of evaluation systems and standards for relevance and value. In addition, the performance of the search engine is different. At present, many search engines focus on the source information on the Internet, ignoring the evaluation of the search engine itself. Therefore, based on existing search engines, a standard system for evaluation and screening of search results should be established for source information.
5ConclusionSince its birth, the search engine has played an important role in Network Information Organization and retrieval. With the explosive growth of information on the Internet in a geometric manner, people's expectations for search engines are getting higher and higher. Although the search engine technology is constantly developing, but at present, especially in China, there are still many deficiencies in the search engine, there is still a long distance from the international advanced level, so it is in line with international practices, according to the actual situation in China, developing a set of search engine standards is an urgent problem.
References1 Fu Shaohong, Huang Miao. research on search engine technology and service and Its Inspiration, Journal of intelligence, 2000 (6) 2 Li Yuanming. analysis of Search Engine Technology and its future development trend, intelligence retrieval, 2002 (7) 3 Wang Hongmei, Zhu Hongxiu, and Wang Ling. A Discussion on the future development of Chinese search engines, Journal of Northeast Electric Power University, (4), Zhang Jun, Chen Yijun. explore the functions and limitations of search engines. intelligence Science, 2001 (5) 5 Tang mingjie. on the development overview and Development Trend of search engines, intelligence magazine, 2001 (5) 6 Yang yingquan, Wen Ru, Huang dengyan. lack of search engines and Application Experience, modern intelligence, 2005 (7) 7 Jose Perez Carballo. natural Language Information Retrival progress report [J]. information Processing and management.8 http://www.sohu. com9 http://www.yahoo. com10 http://www.google. com11 http://www.baidu.com
He ShaohuaProfessor at the School of information management, Wuhan University.
SunChen2005 master of information science, School of information management, Wuhan University.
The content source of this page is from Internet, which doesn't represent Alibaba Cloud's opinion;
products and services mentioned on that page don't have any relationship with Alibaba Cloud. If the
content of the page makes you feel confusing, please write us an email, we will handle the problem
within 5 days after receiving your email.
If you find any instances of plagiarism from the community, please send an email to:
info-contact@alibabacloud.com
and provide relevant evidence. A staff member will contact you within 5 working days.
A Free Trial That Lets You Build Big!
Start building with 50+ products and up to 12 months usage for Elastic Compute Service