Inadvertently see the online Python crawl 1024 of the article, thinking about the late to go to a full-automatic small movie downloader (do not have to choose half a day), work hanging, go back to see (the body has been the sister paper hollowed out, also see), so he first try to write a simple crawler, The goal is naturally the blog park: Use simple regular expression matching, of course, you can also use the widely used online BeautifulSoup parsing web pages
ImportRequestsImportReBASEURL="https://www.cnblogs.com/"HTML=Requests.get (BASEURL). Textitems=re.findall ("_blank\ "> (. +) </a>", HTML) forIinchItems:Print(i)Print("") Print(" Over")
Crawl content is very simple, is the first page of the article list, although C # can also do, but feel python is really streamlined, a few lines of code is done, the word python! The effect is as follows
Can't wait to get back from work and start working!
Python ultra-thin "blog Park" Crawler (sure is more useful than C #)