Python downloads the network novel instance code,
Reading online novels will usually take a wave and import them to the Kindle. However, the mechanical Ctrl + C and Ctrl + V are actually OUT, so this article appears.
In fact, I am also a little white. The main purpose of using Python is its powerful text processing capabilities and network support, as well as many useful libraries. And it's really easier than C (I know it when I use it)
Analyze the webpage to be retrieved
I want to obtain three things:
- The title of the article. <Div id = "title"> Chapter 1 </div>
- Body of the article. <Div id = "content" style = "line-height: 150%; color: rgb (0, 0, 0);">
- The URL of the next chapter. <A href = "11455541.html" rel = "external nofollow"> next page </a>
Also, pay attention to the webpage code. The webpage code is GBK, but in actual operation, I will encounter a webpage decoding error when using GBK:
UnicodeDecodeError: 'gbk' codec can't decode bytes in position 2-3: illegal multibyte sequence
Therefore, after gb18030 is used, the problem will be solved, because in general, there will be a variety of Wang BA's words in xiuxian's network novels. You know, so we need to use a better font, let's take a look at the extensive and profound character encoding.
Source code
I know. You need this, hahaha.
Main Function
# Main function if _ name _ = '_ main __': global numChapter global NOVERL = 'dashboard .txt '# NOVERL = 'tianji .txt 'noverl = Shanghai' if (NOVERL = 'dashboard .txt'): textStartURL = 'HTTP: // your textStartURL = 'HTTP: // www.bxwx8.org/ B /62/62724/28019405.html'## chapter 1,237th ghost master else: textStartURL = 'HTTP: // your textStartURL = 'HTTP: // Combine chapter 78th with the sword technique textStartURL = 'HTTP: // response # textStartURL = 'HTTP: // export nextURL = textStartURL; isEnd = False f = open (NOVERL, 'w', encoding = 'utf-8') f. close () numChapter = 0; while (not isEnd): nextURL, isEnd = findNextTextURL (nextURL) print ('end of capture! ') Print ('obtained' + str (numChapter) + 'zhang ')Get content and URL of the next chapter
# Find the URL of the next chapter # obtain the novel content def findNextTextURL (url): global numChapter global NOVERL # if nextURL = endURL, false if (NOVERL = 'Grand master processor .txt '): endURL = 'HTTP: // your headURL = 'HTTP: // www.bxwx8.org/ B /62/62724/'your master else: endURL = 'HTTP: // your headURL = 'HTTP: // your endURL = 'HTTP: // www.bxwx8.org/ B /35/35282/index.h Tml '# Wudong Qiankun headURL = 'HTTP: // www.bxwx8.org/ B /35/35282/'{wuyunqiankun isEnd = False resp = urllib. request. urlopen (url) # The characters to be processed are indeed gbk characters, but some special characters are included. # It is not included in gbk encoding. If some special characters are included in GB18030, but not in gbk. # Use gbk to decode unsupported characters. For example, an error occurs. # In this case, you can try to decode the code that is compatible with the current encoding (gbk) but contains more characters (gb18030. # AllHtml = resp. read (). decode ('gbk') # allHtml = resp. read (). decode ('gb18030') # textSoup = BeautifulSoup (allHtml) # chapter name strChapter = textSoup. find (id = 'title '). getText (). split (R' [') [0] strChapter = strChapter. split (R' (') [0] strChapter = strChapter. replace ('body', '') + '\ n' numChapter = numChapter + 1 strID =' # '+ str (numChapter) + '-'strchapter = strID + strChapter = strChapter +' \ n character \ n' + url + '\ n ---------------------------- \ n' # novel body strNovel = textSoup. find (id = 'content '). getText () strNovel = strNovel. replace ('', '\ n') # Remove the XXX chapter strMatch = r" section [\ u4e00-\ u9fa5] + chapter "list2replace = re. findall (strMatch, strNovel) if list2replace: str2replace = list2replace [0] strNovel = strNovel. replace (str2replace, '') # merge chapter and body strNovel = strChapter + strNovel + '\ n character \ n ------------------------------ \ n' # Write to write2TXT (strNovel) in the txt file) # obtain the URL nextURL = re. findall (r 'var next_page = "[\ w]+.html" ', allHtml) [0] nextURL = nextURL. replace (R' "','') nextURL = nextURL. replace (r 'var next_page = ', '') nextURL = headURL + nextURL print (numChapter) # number of chapters print (strChapter) # chapter name print (nextURL )) # URL if (endURL = nextURL) in the next chapter: isEnd = True return nextURL, isEndWrite TXT
# Def write2TXT (txt): global NOVERL f = open (NOVERL, 'A', encoding = 'utf-8') f. write (txt + '\ n \ n') f. close ()
Conclusion
Three instructions:
- The orchestration of txt text is definitely not good, and it cannot be automatically divided in the Kindle. You can read more, and the Native system will be GG. Therefore, you can use the epubBuilder software for secondary orchestration, output mobi to your Kindle.
- This program only targets this website, but if the website is changed, the detailed code will have to be rewritten. However, large frameworks can also be used.
- The Internet novels are poisonous to aspiring young people. Once the Internet reaches the Chinese site, it is a sea. From now on, the festival is a passer-by, and the princes are doing what they cherish!
Summary
The above is all the content about downloading the network novel instance code from Python in this article. I hope it will be helpful to you. If you are interested, you can continue to refer to other related topics on this site. If you have any shortcomings, please leave a message. Thank you for your support!