Simple python collaborative filtering program instance code and python collaborative filtering instance
This article focuses on the content of the python collaborative filtering program.
One of the most typical examples of collaborative filtering is watching movies. Sometimes we don't know which movie is our favorite or has a high rating. The common practice is to ask friends around us, let's take a look at some good movie recommendations recently. When you ask, you are used to asking friends who have similar tastes. This is the core idea of collaborative filtering.
This program is a small program written to cope with the big data analysis and Computing Course assignments, first on the program, a total of 55 lines. If you don't care about the details, the 55-line program has shown the feature of collaborative filtering. It is to find four closest users for each user and then make recommendations. When selecting recommendations, we directly select items not included in the four users, of course, there is no limit on the number of recommendations. I personally think that if you want to improve the accuracy of recommendations, at least, 1, you need to process popular items. 2. Sort the items of the four adjacent users, and perform recommendation from many to few. The data used by the program is movielens (http://grouplens.org/datasets/movielens ). Similarity calculation is also very simple, and the ratio of intersection to difference set is used directly. Okay, go to the program
# Coding utf-8import osimport sysimport ref1 = open ("/home/alber/data_base/bigdata/movielens_train_result.txt", 'R') # Read the train file, each row has been processed to represent a user's item, and spaces are used between items. F2 = open ("/home/alber/data_base/bigdata/movielens_train_result3.txt", 'A') txt = f1.readlines () contxt = [] f1.close () userdic ={} for line in txt: line_clean = "". join (line. split () position = line_clean.index (",") ID = line_clean [0: position] item = line_clean [position + 1:] userdic. setdefault (ID, item) if len (item)> = 5: # users with less than 5 views are not included in the contxt range of similarity calculation. append (item) for key in userdic. keys (): # Calculate the four most similar users of each user ID_num = key valu E = userdic [key] user_item = value. split ('') Sim_user = [] for lines in contxt: lines_clean = lines. split ('') intersection = list (set (lines_clean ). intersection (set (user_item) lenth_intersection = len (intersection) difference = list (set (lines_clean ). difference (set (user_item) lenth_difference = len (difference) if lenth_difference! = 0: Similarity = float (lenth_intersection)/lenth_difference # divide the intersection by the difference set as the Similarity judgment condition (Similarity) else: Sim_user.append ("0") Sim_user_copy = Sim_user [:] round () Sim_best = Sim_user_copy [-4:] position1 = Sim_user.index (Sim_best [3]) position2 = Sim_user.index (Sim_best [2]) position3 = Sim_user.index (Sim_best [1]) position4 = Sim_user.index (Sim_best [0]) if position1! = 0 and position2! = 0 and position3! = 0 and position4! = 0: recommender = userdic [str (position1)] + "" + userdic [str (position2)] + "" + userdic [str (position3)] + "" + userdic [str (position4)] # recommended else: recommender = "none" reco_list = recommender. split ('') recomm = [] for good in reco_list: if good not in user_item: recomm. append (good) else: pass f2.write (("". join (recomm) + "\ n") f2.close ()
Summary
The above is all the content about the simple python collaborative filtering program instance code. I hope it will be helpful to you. If you are interested, you can continue to refer to other related topics on this site. If you have any shortcomings, please leave a message. Thank you for your support!