Javascript| Search Engine | page
In the production of search engines, or do page analysis and data extraction, often faced with a lot of JavaScript in the page, and the page content, a considerable part of the written to these JS script commands, and resulting in normal DOM analysis failed to extract the required information. Of course, if this page template to determine, for this page to make information extraction template is not difficult, each page manually analyze the need to extract the location of information, and then make a template. But for general's web search, this is not realistic. It happened two days ago to discuss the problem with friends, some ideas. Here, provide two ideas, for your reference. 1, a simplified JavaScript interpreter, the execution of script fragments to do a complete JavaScript interpreter is relatively rare, but it is easy to make a simplified JavaScript interpreter. We don't need those complex libraries, we just implement the basic JavaScript syntax and implement part of the function that involves text output. The goal is not to actually execute the JavaScript completely, but to combine the strings in the script, according to their program logic, and finally output the full output of the script. This nature is not comprehensive, certainly because many functions are not implemented, resulting in the output string and the real output is not exactly the same. But, if it's not an accident,
When making a search engine, or doing page analysis and data extraction, often face a lot of JavaScript on the page, these javascript is annoying, because a considerable part of the page content written to these JS script commands, and cause normal DOM analysis can not see these words , which causes the text data to be extracted failed.
Of course, if this page template to determine, for this particular page to make information extraction template is not difficult, each page manually analyze the need to extract the location of information, and then make a template. But for general's web search, this is not realistic. It happened two days ago to discuss the problem with friends, some ideas. Here, provide two ideas, for your reference.
1. Make a simplified JavaScript interpreter, execute script fragment
Making a complete JavaScript interpreter is rare, but making a simplified JavaScript interpreter is easy. We don't need those complex libraries, we just implement the basic JavaScript syntax and implement part of the function that involves text output.
The goal is not to actually execute the JavaScript completely, but to combine the strings in the script, according to their program logic, and finally output the full output of the script. This nature is not comprehensive, certainly because many functions are not implemented, resulting in the output string and the real output is not exactly the same. However, if there is no accident, it should not produce too many omissions. Because all of the string output parts are implemented, it is perfectly possible to combine the strings with the logic that they will output.
For things that are dynamic based on dynamic conditions, if these conditions are not certain, for example, depending on the browser type or whatever. The results of all two branches can be completely exported. Of course, we should not combine these two pieces of text, the middle should have our understanding of the separator.
The advantage of doing this is high performance. This interpreter can be done very small, because it is not the full implementation of JS, so the performance is also faster. The disadvantage is that because it is a simplified interpreter, there is a difference between the actual result and the real one. But in general, the information will only be more and not less, (because of the output of different branches of the results), so for the search engine page analysis, almost enough.
2, with HTML rendering engine complete parsing page, and finally from the display results to take data
Use Gecko (Firefox) or Trident (Mshtml.dll) (IE) HTML rendering engine for browsers to complete parsing and rendering of the page. Finally, the analytic results of these engines are analyzed.
The advantage of doing this is that it is closest to the display result, because they are the actual results of the page parsing. But the disadvantage is that the performance is relatively poor, because it is a complete analysis of all elements of the page, so do a lot of work with the extraction of text information useless labor, if the analysis of large amount of data page, need to weigh.