Preface
The script in this paper is to analyze the access log of Nginx, mainly to check the number of access points URI, the results will be provided to developers for reference, because when it comes to analysis, it is necessary to use regular expression, so please do not contact the regular partner self-brain, because of the content of the regular, It is impossible to write, the content of the regular is too large, is not an article two can be written clearly.
Before we start, let's look at the log structure to analyze:
127.0.0.1--[19/jun/2012:09:16:22 +0100] "get/go.jpg http/1.1" 499 0 "http://domain.com/htm_data/7/1206/758536.html" " mozilla/4.0 (compatible; MSIE 7.0; Windows NT 5.1; trident/4.0;. NET CLR 1.1.4322;. NET CLR 2.0.50727;. NET CLR 3.0.4506.2152;. NET CLR 3.5.30729; SE 2.X METASR 1.0) "127.0.0.1--[19/jun/2012:09:16:25 +0100]" get/zyb.gif http/1.1 "499 0" http://domain.com/htm_data/7/ 1206/758536.html "" mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1; Qqdownload 711; SV1;. net4.0c;. net4.0e; 360SE) "
This is the modified log content, the sensitive content is deleted or replaced, but does not affect our analysis results, of course, the format of what is not important, nginx access logs can be customized, each company may be slightly different, so to understand the script content, And through their own changes applied to their own work is the focus, I give the log format is a reference, I bet you see on your company server log format is certainly not the same as my format, read the log format, we began to write our script
I'll post the code and explain later:
Import refrom operator Import Itemgetter def parser_logfile (logfile): pattern = (r "' (\d+.\d+.\d+.\d+) \s-\s-\s ' #IP Address ' \[(. +) \]\s ' #datetime ' "get\s (. +) \s\w+/.+" \s ' #requested file ' (\d+) \s ' #status ' (\d+) \s ' #bandwidth ' "(. +)" \s ' #referrer ' "(. +)" ' #user agent ' ) fi = open (logfile, ' r ') url_list = [] for line in fi:
url_list.append (Re.findall (pattern, line)) Fi.close () return url_list def parser_urllist (url_list): URLs = [] for URL I N url_list: for R in URL: urls.append (r[5]) return URLs def get_urldict (URLs): D = {} for URL in URLs: d[url ] = D.get (url,0) +1 return D def url_count (logfile): Url_list = parser_logfile (logfile) urls = parser_urllist (url_list) tot als = get_urldict (URLs) return totals if __name__ = = ' __main__ ': urls_with_counts = Url_count (' Example.log ') Sorted_by_cou NT = sorted (Urls_with_counts.items (), Key=itemgetter (1), reverse=true) print (Sorted_by_count)
Script interpretation, parser_logfile() function function is to analyze the log, return the matching list of rows, the regular part does not explain, we see the note should know it is matching what content, parser_urllist() function function is to get the URL of the user access, get_urldict() function function is to return a dictionary, the URL is the key, If the same value of the key is increased by 1, the returned dictionary is each URL and the maximum number of accesses, url_count() function function is to call the previously defined function, the main function part, say itemgetter, it can be implemented by the specified elements of the order, for example, understand:
>>> from operator import itemgetter>>> a=[(' B ', 2), (' A ', 1), (' C ', 0)] >>> s=sorted (a,key= Itemgetter (1)) >>> s[(' C ', 0), (' A ', 1), (' B ', 2)]>>> s=sorted (a,key=itemgetter (0)) >>> s[(' A ', 1), (' B ', 2), (' C ', 0)]
The Reverse=true parameter indicates descending sort, that is, from large to small, the script runs the result:
[(' http://domain.com/htm_data/7/1206/758536.html ', 141), (' http://domain.com/?q=node&page=12 ', 3), (' HTTP// Website.net/htm_data/7/1206/758536.html ', 1)]