How does the engine index webpages?

Source: Internet
Author: User

Highdiy was published in May 9, 2007
For SEO (Search Engine Optimization), it should be said that it is the primary task to enable the pages on the website to be indexed and indexed in a timely and comprehensive manner by the search engine. This is the most basic guarantee for implementing other SEO policies. -- However, this is also a process that is often overestimated, for example, we often see how many pages, such as several K or even dozens of K, are claimed to be indexed by Google on some websites to prove the success of SEO. But objectively speaking, simply indexing and indexing web pages by search engines does not have much practical significance. It can only be used as a funeral product in the vast Internet world, more importantly, how to make a webpage appear in the first few pages of the SERP (search result page) for a specific search item. -- Many people believe that it is not a bad thing to include as many pages as possible in the website into the search engine index database. The more webpages, the greater the chance of exposure, although there are doubts about the final effect.
Anyway, if you focus on the speed and efficiency of Web Page indexing and indexing when implementing SEO for your website, it is understandable. To achieve this, we need to have some knowledge about how search engines include and index Web pages. Next we will take Google as an example to introduce the process of indexing and indexing web pages by search engines, hoping to help our friends later. -- For other search engines such as Yahoo! In terms of Live Search and Baidu, although there may be differences in specific details, the basic strategy should be similar.
1. Collect the url of the webpage to be indexed
The number of web pages on the Internet is definitely an astronomical number, and there are countless new web pages each day. Search Engines need to first find the objects to be indexed.
Specifically for Google, although there is a dispute over whether there is a difference between DeepBot and FreshBot in GoogleBot-whether such two names are called is even more difficult, of course, the name itself is not important-at least so far, the mainstream view is that in Google's robots, there are indeed a considerable number of robots dedicated to preparing "Materials" for the real index pages -- here we will call it FreshBot -- their task is to constantly scan the Internet every day, to discover and maintain a large list of URLs for DeepBot. In other words, when it accesses and reads a webpage, it does not aim to index the webpage, but to find all the links in the webpage. -- Of course, there seems to be a conflict in efficiency, which is a bit untrustworthy. However, we can simply judge by using the following method: FreshBot does not have "ranking" when scanning a webpage, that is to say, multiple robots located in different data centers of Google may access the same page within a short period of time, for example, one day or even one hour, however, DeepBot does not perform similar operations when indexing or caching pages, that is, Google will restrict robots from a data center to complete this task, the two data centers will not index the same version of the web page at the same time. If there is no flaw in this statement, from the server access log, it seems that the GoogleBot from different IP addresses often accesses the same web page multiple times in a short time to prove the existence of FreshBot. Therefore, you may not be too happy to find that GoogleBot frequently accesses the website. Maybe it is not indexing the webpage, but scanning the url.
The information recorded by FreshBot includes the webpage url, Time Stamp (the timestamp when the webpage is created or updated), and the webpage Head information (Note: This is controversial, there are also many people who believe that FreshBot will not read the target webpage information, but will hand over this work to DeepBot. A considerable portion of the external tables are implemented through the "noindex" in the mata tag. It seems that this cannot be achieved without reading the head of the target webpage). If the webpage is inaccessible, for example, if the network is interrupted or the server is faulty, FreshBot will write down the url and choose to retry. However, it will not be added to the list of URLs submitted to DeepBot before the url is accessible.
In general, FreshBot still occupies a relatively small amount of server bandwidth and resources. Finally, FreshBot classifies the record information according to different priorities and submits the information to DeepBot. There are mainly the following types based on different priorities:
A: create A webpage;
B: The old webpage/new Time Stamp, that is, the updated webpage;
C: Use a 301/302 redirected webpage;
D: complex dynamic URLs. If multiple parameters are used for dynamic URLs, Google may need additional work to analyze the content correctly. -- With the improvement of Google's support for dynamic web pages, this category may have been canceled;
E: other types of files, such as links to PDF and DOC files, may need additional work for indexing these files;
F: The old webpage or the old Time Stamp, that is, the unupdated webpage. Note that the timestamp here is not based on the date displayed in the Google search results, it is compared with the date in the Google index database;
G: incorrect url, that is, the 404 response page is returned during access;
Priorities are arranged in the order of A to G. It should be emphasized that the priority mentioned here is relative. For example, when a new webpage is created, the priority varies greatly depending on the link quality and quantity, websites with links from relevant authoritative websites have a high priority. In addition, the priority here is only for pages within the same website. In fact, different websites have different priorities. In other words, for webpages on authoritative websites, even the 404 url with the lowest priority may be more advantageous than creating a webpage with the highest priority for many other websites.
2. webpage indexing and Indexing
The next step is to go to the real indexing and indexing web page process. From the above introduction, we can see that the list of URLs submitted by FreshBot is quite large. According to different languages and website locations, indexing of specific websites will be distributed to different data centers. The entire indexing process may take weeks or even longer to complete due to the large amount of data.
As mentioned above, DeepBot will first index websites/WebPages with a higher priority. The higher the priority, the faster it will appear on the Google index database and eventually on the Google search results page. As long as you enter this stage for creating a web page, even if the entire indexing process is not completed, the corresponding web page may already appear in the Google index library, I believe many of my friends often see pages marked as "site: somedomain.com" in Google that show only webpage URLs or only webpage titles and URLs but not descriptions, this is the normal result of the webpage at this stage. After Google actually reads, analyzes, and caches this page, it will escape from the supplementary results and display normal information. -- Of course, the premise is that the webpage has enough links, especially links from authoritative websites, and the index library does not have the same or similar records as the webpage Content (Duplicate Content filtering ).
For dynamic URLs, although Google claims that there is no obstacle in its processing, it can be observed that the probability of dynamic URLs appearing in the supplementary results is much higher than that of webpages using static URLs. More and more valuable links are often needed to escape from the supplementary results.
For the above "F" class, that is, the unupdated webpage, DeepBot compares its timestamp with the date in the Google index database, make sure that, although the corresponding page information in the search results may not be updated in time, you only need to index the latest version-consider the situation where the web page is updated and modified multiple times-as for the "G" class, that is, 404 url, the system checks whether there are corresponding records in the index database. If yes, it deletes the records.
3. Data Center Synchronization
As mentioned above, DeepBot indexes a webpage by a specific data center, instead of reading the webpage from multiple data centers and obtaining the latest version of the webpage, after the indexing process is complete, a data synchronization process is required to update the latest version of the web page in multiple data centers.
This is the famous Google Dance. However, after the BigDaddy update, synchronization between data centers is no longer concentrated in a specific period of time, but in a continuous and time-sensitive manner. Although there are still some differences between different data centers, the difference is not big, and the maintenance time is very short.
To improve the efficiency of indexing web pages by search engines, we can see that you should optimize your web pages to be indexed as quickly and as many as possible by search engines:
Improve the quantity and quality of reverse link for your website. links from authoritative websites can be "seen" by the search engine immediately ". Of course, this is also a commonplace. From the above introduction, we can see that to improve the efficiency of webpage indexing by search engines, we must first let the search engine Find Your webpage, links are the only way for search engines to find web pages-the word "unique" is somewhat controversial. See the SiteMaps section below-from this perspective, submitting a website to a search engine is unnecessary and meaningless. To include your website, obtaining a link to an external website is fundamental. At the same time, high-quality links are also a key factor in making webpages more productive.
Webpage Design should adhere to the "Friendly search engine" principle, design and optimize the webpage from the perspective of search engine spider, and ensure that the internal links of the website are "visible" to the search engine ", compared with the difficulty of obtaining external website links, a reasonably planned internal link is a more economical and effective way to improve search engine indexing and indexing efficiency-unless the website is not indexed by a search engine at all.
If your website uses dynamic URLs or the navigation menu uses JavaScript, you should first start from here when encountering obstacles in Webpage indexing.
Use SiteMaps. In fact, many people think that one of the main reasons Google has canceled FreshBot is the wide application of SiteMaps (xml) Protocol. In this way, you only need to read the SiteMaps provided by the website to obtain the webpage update information, freshBot does not require time-consuming and laborious scanning. This statement is still true. Although it is difficult to determine whether Google uses SiteMaps directly as the index list of DeepBot or uses FreshBot as the scanning roadmap, however, SiteMaps can improve the indexing efficiency of websites. For example, SEO has been tested as follows:
The links obtained from the two web pages are the same. When one is added to SiteMaps and the other is not added, the pages that appear in SiteMaps will soon be included, the other page is indexed after a long time;
There is no link to an island page, but after adding it to SiteMaps for a while, it is also indexed by Google, but it appears in the supplementary result.
Of course, although the webpage is not found in SiteMaps, it can still be indexed by Google. Google still uses FreshBot or a mechanism similar to FreshBot, which is easy to understand, after all, there are still so many websites that are not using SiteMaps that Google cannot reject.
For more information about SiteMaps, see Google SiteMaps: Google's "backdoor ". It should be noted that the Sitemaps Protocol has become an industry standard and is not only effective for Google. Other mainstream search engines include Yahoo! , Live search, and Ask are supported.
Disclaimer: part of the information in this article is from the public literature, and part is purely a personal speculation. You may be surprised to hear this.
Author:
Highdiy
Original: dianshi Interaction
Search Engine Optimization
Blog
Copyright Disclaimer: This article has been authorized by the author to publish. Please keep this copyright information for reprinting and prohibit unauthorized copying.

Contact Us

The content source of this page is from Internet, which doesn't represent Alibaba Cloud's opinion; products and services mentioned on that page don't have any relationship with Alibaba Cloud. If the content of the page makes you feel confusing, please write us an email, we will handle the problem within 5 days after receiving your email.

If you find any instances of plagiarism from the community, please send an email to: info-contact@alibabacloud.com and provide relevant evidence. A staff member will contact you within 5 working days.

A Free Trial That Lets You Build Big!

Start building with 50+ products and up to 12 months usage for Elastic Compute Service

  • Sales Support

    1 on 1 presale consultation

  • After-Sales Support

    24/7 Technical Support 6 Free Tickets per Quarter Faster Response

  • Alibaba Cloud offers highly flexible support services tailored to meet your exact needs.