This article mainly introduces how to use Python to compile the BT resource crawler based on the DHT protocol. This article also introduces the knowledge about the DHT protocol. For more information, see
About DHT Protocol
As an aid to the BT protocol, the DHT protocol is very interesting. It is mainly used to get the seeds or BT resources when BT is officially downloaded. Traditional networks require a central server to store seeds or BT resources, which is not only a waste of server resources, but also prone to spof issues. The DHT network is for decentralization, that is, anytime, there are always bright nodes in this network. You can ask these bright nodes to add yourself to the DHT network.
To implement the network crawler of the DHT protocol, the first step is to obtain the resource information (infohash, 160bit, 20 bytes, can be encoded as a 40-byte hexadecimal string ), the second step is to confirm that these infohashes are valid, and the third step is to download the valid infohash to the BT seed file to obtain a complete description of this resource.
The first step is that other nodes use the get_peers method in the DHT protocol to send requests to crawlers. The second step is that other nodes use the announce_peer method in the DHT protocol to send requests to crawlers, step 3 can be obtained in several ways. For example, you can go to some websites that store seeds to download them directly based on infohash, or download them through announce_peer nodes, it depends on your crawlers.
The main operations in the DHT protocol are as follows:
It is mainly responsible for interacting with external nodes through UDP, encapsulating 4 kinds of basic operation requests and corresponding.
Ping: Check whether a node is "alive"
Ping is mainly used in two aspects of a crawler. The first is the initial route table, and the second is to verify whether the node is alive.
Find_node: Send a node search request to a node
In a crawler, find_node is mainly used in two places. The first is the initial route table, and the second is to verify whether the bucket is alive.
Get_peers: sends a resource search request to a node.
When a crawler has a node to send a request to itself, it not only responds like a normal node, but also needs to take the info_hash of this resource as an opportunity to know as many nodes as possible ., The last step of get_peers is actually announce_peer, but because crawlers cannot announce_peer, get_peers actually degrades to the find_node operation.
Announce_peer: sends a notification to a node that you have started to download a resource.
Announce_peer cannot be used in crawlers, because it is equivalent to reporting false resources, and the other party can easily determine from the context whether you have reported false resources to ban you.
Python-based DHT Crawler
Modified from github open-source crawler. The original author's name is somewhat... here, the Project address is listed directly: Ghost
#!/usr/bin/env python# encoding: utf-8import socketfrom hashlib import sha1from random import randintfrom struct import unpackfrom socket import inet_ntoafrom threading import Timer, Threadfrom time import sleepfrom collections import dequefrom bencode import bencode, bdecodeBOOTSTRAP_NODES = ( ("router.bittorrent.com", 6881), ("dht.transmissionbt.com", 6881), ("router.utorrent.com", 6881))TID_LENGTH = 2RE_JOIN_DHT_INTERVAL = 3TOKEN_LENGTH = 2def entropy(length): return "".join(chr(randint(0, 255)) for _ in xrange(length))def random_id(): h = sha1() h.update(entropy(20)) return h.digest()def decode_nodes(nodes): n = [] length = len(nodes) if (length % 26) != 0: return n for i in range(0, length, 26): nid = nodes[i:i+20] ip = inet_ntoa(nodes[i+20:i+24]) port = unpack("!H", nodes[i+24:i+26])[0] n.append((nid, ip, port)) return ndef timer(t, f): Timer(t, f).start()def get_neighbor(target, nid, end=10): return target[:end]+nid[end:]class KNode(object): def __init__(self, nid, ip, port): self.nid = nid self.ip = ip self.port = portclass DHTClient(Thread): def __init__(self, max_node_qsize): Thread.__init__(self) self.setDaemon(True) self.max_node_qsize = max_node_qsize self.nid = random_id() self.nodes = deque(maxlen=max_node_qsize) def send_krpc(self, msg, address): try: self.ufd.sendto(bencode(msg), address) except Exception: pass def send_find_node(self, address, nid=None): nid = get_neighbor(nid, self.nid) if nid else self.nid tid = entropy(TID_LENGTH) msg = { "t": tid, "y": "q", "q": "find_node", "a": { "id": nid, "target": random_id() } } self.send_krpc(msg, address) def join_DHT(self): for address in BOOTSTRAP_NODES: self.send_find_node(address) def re_join_DHT(self): if len(self.nodes) == 0: self.join_DHT() timer(RE_JOIN_DHT_INTERVAL, self.re_join_DHT) def auto_send_find_node(self): wait = 1.0 / self.max_node_qsize while True: try: node = self.nodes.popleft() self.send_find_node((node.ip, node.port), node.nid) except IndexError: pass sleep(wait) def process_find_node_response(self, msg, address): nodes = decode_nodes(msg["r"]["nodes"]) for node in nodes: (nid, ip, port) = node if len(nid) != 20: continue if ip == self.bind_ip: continue if port < 1 or port > 65535: continue n = KNode(nid, ip, port) self.nodes.append(n)class DHTServer(DHTClient): def __init__(self, master, bind_ip, bind_port, max_node_qsize): DHTClient.__init__(self, max_node_qsize) self.master = master self.bind_ip = bind_ip self.bind_port = bind_port self.process_request_actions = { "get_peers": self.on_get_peers_request, "announce_peer": self.on_announce_peer_request, } self.ufd = socket.socket(socket.AF_INET, socket.SOCK_DGRAM, socket.IPPROTO_UDP) self.ufd.bind((self.bind_ip, self.bind_port)) timer(RE_JOIN_DHT_INTERVAL, self.re_join_DHT) def run(self): self.re_join_DHT() while True: try: (data, address) = self.ufd.recvfrom(65536) msg = bdecode(data) self.on_message(msg, address) except Exception: pass def on_message(self, msg, address): try: if msg["y"] == "r": if msg["r"].has_key("nodes"): self.process_find_node_response(msg, address) elif msg["y"] == "q": try: self.process_request_actions[msg["q"]](msg, address) except KeyError: self.play_dead(msg, address) except KeyError: pass def on_get_peers_request(self, msg, address): try: infohash = msg["a"]["info_hash"] tid = msg["t"] nid = msg["a"]["id"] token = infohash[:TOKEN_LENGTH] msg = { "t": tid, "y": "r", "r": { "id": get_neighbor(infohash, self.nid), "nodes": "", "token": token } } self.send_krpc(msg, address) except KeyError: pass def on_announce_peer_request(self, msg, address): try: infohash = msg["a"]["info_hash"] #print msg["a"] tname = msg["a"]["name"] token = msg["a"]["token"] nid = msg["a"]["id"] tid = msg["t"] if infohash[:TOKEN_LENGTH] == token: if msg["a"].has_key("implied_port") and msg["a"]["implied_port"] != 0: port = address[1] else: port = msg["a"]["port"] if port < 1 or port > 65535: return self.master.log(infohash, (address[0], port),tname) except Exception: pass finally: self.ok(msg, address) def play_dead(self, msg, address): try: tid = msg["t"] msg = { "t": tid, "y": "e", "e": [202, "Server Error"] } self.send_krpc(msg, address) except KeyError: pass def ok(self, msg, address): try: tid = msg["t"] nid = msg["a"]["id"] msg = { "t": tid, "y": "r", "r": { "id": get_neighbor(nid, self.nid) } } self.send_krpc(msg, address) except KeyError: passclass Master(object): def log(self, infohash,address=None,tname=None): hexinfohash = infohash.encode("hex") print "info_hash is: %s,name is: %s from %s:%s" % ( hexinfohash,tname, address[0], address[1] ) print "magnet:?xt=urn:btih:%s&dn=%s" % (hexinfohash, tname)# using exampleif __name__ == "__main__": # max_node_qsize bigger, bandwith bigger, speed higher dht = DHTServer(Master(), "0.0.0.0", 6882, max_node_qsize=200) dht.start() dht.auto_send_find_node()
PS: the DHT protocol has several key points to be clarified:
1. node and infohash use the same 160-bit representation. The 730750818665451459101842416358141509827966271488-bit representation means that the entire node space is 2 ^ =, which is in a 48-bit 10-hexadecimal format. That is to say, there are billion and million node spaces, such a large Node space is sufficient to store your host nodes and any resource information.
2. Each node has a route table. Each route table consists of a bunch of K buckets. The so-called K-bucket means that a maximum of K nodes can be placed in the bucket. The default value is 8. Bucket storage is similar to a Prefix Tree. It is equivalent to a maximum of to 4 K buckets in an 8-bucket route table.
3. according to the DHT Protocol, each infohash has a location. Therefore, there is a distance between two infohash, and the distance between the two infohash can be expressed by an exclusive or, that is, infohash1 xor infohash2. That is to say, if the height is the same, the distance between them is the nearest. If the height is the same, the distance between the two nodes can be quickly calculated. What is the purpose of calculating this distance? In the DHT network, if the infohash of a resource is closer to the infohash of a node, the node is more likely to have information about the resource. Why? As you can imagine, everyone uses the same distance algorithm to recursively ask the node close to the resource, and as long as the node responds, it will get an announce information, that is to say, nodes close to the resource infohash have a higher probability to get the infohash of the resource.
4. Based on the above algorithm, the query in the DHT is a skip query, which can quickly bridge the node bucket to the target node bucket. The reason why a large jump can be made in the distance, but only a small jump can be made in the near future, is that the closer each node's routing table is to itself, the more it stores, as shown in
5. in a DHT network, crawlers are not easy to crawl. Unlike Common crawlers, You can automatically crawl resources when you see them. On the contrary, the method for obtaining resources (get_peers, announce_peer) is passive, the crawler method has changed. The crawler must respond to queries from other nodes like a normal node and receive responses from other nodes, collecting the data is a task. The only thing that crawlers can do is to try to recognize other nodes as much as possible so that more nodes can ask you.
6. some people have said that if I increase the capacity K in the K bucket of the DHT crawler, will it be able to increase the chance of obtaining resources? Actually, I have analyzed it before, the most important information sources of DHT crawlers are passive, because you cannot increase others' K, the smaller the probability of saving yourself to a distant node, of course, the probability to request you from a distant node is relatively small.