Using Python to write a BT resource crawler based on the DHT protocol

Source: Internet
Author: User
about the DHT protocol

The DHT protocol as an aid to the BT protocol is very fun. It is primarily intended to get seeds or BT resources when BT is officially downloaded. Traditional network, need a central server to store seeds or BT resources, not only waste server resources, but also prone to a single point of various problems, and the DHT network is to be centralized, that is, at any time, the network always has a node is bright, you can ask these bright nodes, Thereby adding itself to the DHT network.

To implement the DHT protocol of the web crawler, mainly in 3 steps, the first step is to get the resource information (infohash,160bit,20 bytes, can be encoded into a 40-byte hexadecimal string), the second step is to confirm that these infohash are valid, The third step is to get a complete description of this resource by downloading the torrent file to BT via a valid Infohash.

The first step is that other nodes use the Get_peers method in the DHT protocol to send requests to the crawler, the second step is that the other nodes with the DHT protocol Announce_peer to the crawler to get the request, the third step can be obtained in several ways, For example, you can go to some of the preservation of the seed site according to Infohash direct download to, or through the Announce_peer node to download to, how to achieve, depending on your own crawler.

The main operations in the DHT protocol are:

It is primarily responsible for interacting with external nodes via UDP, encapsulating requests for 4 basic operations and corresponding.

Ping: Check if a node is "alive"

In a reptile there are two main places to use Ping, the first is the initial routing table, the second is to verify that the node is alive

Find_node: A request to send a lookup node to a node

In a reptile, the main is two places to use Find_node, the first is the initial routing table, the second is to verify that the bucket is alive

Get_peers: Sending a request to find a resource to a node

There are nodes in the crawler that respond to their requests not only as a normal node, but also with the info_hash of this resource as much as possible to get to know more nodes. , Get_peers actually the last step is announce_peer, but because the crawler can not announce_peer, so actually get_peers degenerate into find_node operation.

Announce_peer: Sends a notification to a node that it has started downloading a resource

Reptiles can not use announce_peer, because this is equivalent to the notification of false resources, the other side easily from the context to determine whether you have informed the false resources to ban you.

Python-based DHT crawler
Modified from GitHub Open source Crawler, the original author name some. , the project address is listed directly here: Https://github.com/Fuck-You-GFW/simDHT, who has a GitHub account with the original author star, and later I put the results into DB, Plus use Tornado to make a simple query interface and put it on GitHub, back up the code first.

#!/usr/bin/env python# encoding:utf-8import socketfrom hashlib import sha1from random import randintfrom struct import un  Packfrom Socket Import inet_ntoafrom Threading import Timer, Threadfrom time import sleepfrom collections Import Dequefrom  Bencode import Bencode, Bdecodebootstrap_nodes = (("router.bittorrent.com", 6881), ("dht.transmissionbt.com", 6881), ("router.utorrent.com", 6881)) Tid_length = 2re_join_dht_interval = 3token_length = 2def Entropy (LENGTH): Return "". JOIN (Chr (randint (0, 255)) for _ in X Range (length) def random_id (): H = SHA1 () h.update (entropy ()) return H.digest () def decode_nodes (nodes): n = [] Leng th = Len (nodes) if (length%)! = 0:return N for i in range (0, length, n): nid = nodes[i:i+20] IP = inet_nt OA (nodes[i+20:i+24]) port = Unpack ("! H ", nodes[i+24:i+26]) [0] N.append (NID, IP, Port)) return Ndef timer (T, f): Timer (T, f). Start () def Get_neighbor (targe T, nid, End=10): Return Target[:end]+nid[end:]class KNode (object): Def __iNit__ (self, NID, IP, port): Self.nid = nid self.ip = IP Self.port = portclass dhtclient (Thread): def __init__ (SE LF, max_node_qsize): thread.__init__ (self) self.setdaemon (True) self.max_node_qsize = Max_node_qsize Self.nid = random_id () self.nodes = Deque (maxlen=max_node_qsize) def send_krpc (self, MSG, address): Try:self.ufd.sendt O (Bencode (msg), address) except Exception:pass def send_find_node (self, Address, nid=none): nid = Get_neighbo R (Nid, Self.nid) if nid else Self.nid tid = entropy (tid_length) msg = {"T": tid, "y": "Q", "Q": "Fin D_node "," a ": {" id ": Nid," target ": random_id ()}} SELF.SEND_KRPC (msg, address) def Join_ DHT (self): to address in BOOTSTRAP_NODES:self.send_find_node (address) def re_join_dht (self): If Len (Self.nod ES) = = 0:SELF.JOIN_DHT () timer (Re_join_dht_interval, SELF.RE_JOIN_DHT) def auto_send_find_node (self): wait = 1.0/self.max_node_qsizE while True:try:node = Self.nodes.popleft () self.send_find_node ((Node.ip, Node.port), Node.nid) Except Indexerror:pass sleep (wait) def process_find_node_response (self, MSG, address): nodes = Decode  _nodes (msg["R" ["Nodes"]) for node in nodes: (Nid, IP, port) = node If Len (nid)! = 20:continue if IP = = Self.bind_ip:continue if port < 1 or port > 65535:continue n = KNode (nid, IP, port) self.nodes.app End (N) class Dhtserver (dhtclient): Def __init__ (self, master, Bind_ip, Bind_port, max_node_qsize): dhtclient.__init__ (S Elf, max_node_qsize) Self.master = Master Self.bind_ip = bind_ip Self.bind_port = Bind_port Self.process_reque St_actions = {"Get_peers": Self.on_get_peers_request, "Announce_peer": self.on_announce_peer_request,} s ELF.UFD = Socket.socket (socket.af_inet, socket. SOCK_DGRAM, Socket. IPPROTO_UDP) Self.ufd.bind ((Self.bind_ip, Self.bind_port)) timer (Re_join_dHt_interval, SELF.RE_JOIN_DHT) def run (self): SELF.RE_JOIN_DHT () while True:try: (data, address) = Sel  F.ufd.recvfrom (65536) msg = bdecode (data) Self.on_message (MSG, address) except Exception:pass def on_message (self, MSG, address): Try:if msg["y"] = = "R": if Msg["R"].has_key ("Nodes"): SELF.PR Ocess_find_node_response (msg, address) elif msg["y"] = = "Q": try:self.process_request_actions[msg["Q "]] (MSG, address) except KeyError:self.play_dead (MSG, address) except Keyerror:pass def on_get_ Peers_request (Self, MSG, address): Try:infohash = msg["a" ["info_hash"] tid = msg["T"] nid = msg["a" [" ID "] token = infohash[:token_length] msg = {" T ": tid," y ":" R "," R ": {" id ": get_ Neighbor (Infohash, Self.nid), "nodes": "", "token": Token}} self.send_krpc (MSG, Addre SS) except Keyerror:     Pass Def on_announce_peer_request (self, MSG, address): Try:infohash = msg["a" ["Info_hash"] #print msg ["a"] Tname = msg["A" ["name"] token = msg["a" ["token"] nid = msg["a" ["id"] tid = msg["T"] if I Nfohash[:token_length] = = Token:if msg["A"].has_key ("Implied_port") and msg["a" ["implied_port"]! = 0:po RT = Address[1] Else:port = msg["A" ["Port"] if port < 1 or port > 65535:return SE Lf.master.log (Infohash, (Address[0], port), tname) except Exception:pass Finally:self.ok (MSG, address) d EF Play_dead (Self, MSG, address): Try:tid = msg["T"] msg = {"T": tid, "y": "E", "E": [    202, "Server Error"} self.send_krpc (msg, address) except Keyerror:pass def ok (self, MSG, address):          Try:tid = msg["T"] nid = msg["a" ["id"] msg = {"T": tid, "y": "R", "R": { "id": Get_neighbor (Nid, Self.nid)}} SELF.SEND_KRPC (msg, address) except Keyerror:passclass Master (object): Def log (s Elf, infohash,address=none,tname=none): Hexinfohash = Infohash.encode ("hex") print "Info_hash is:%s,name is:%s fro  M%s:%s "% (Hexinfohash,tname, address[0], address[1]) print" magnet:?xt=urn:btih:%s&dn=%s "% (Hexinfohash, Tname) # using ExampleIf __name__ = = "__main__": # max_node_qsize bigger, bandwith bigger, speed higher DHT = Dhtserver ( Master (), "0.0.0.0", 6882, max_node_qsize=200) Dht.start () Dht.auto_send_find_node ()

The PS:DHT agreement has several areas of focus that need clarification:

1. Node and Infohash also use the 160bit representation, 160bit means the entire node space has 2^160 = 730750818665451459101842416358141509827966271488, is 48 bit 10 binary , that is, there are exascale billion of node space, so large node space, is enough to store your host node and any resource information.

2. Each node has a routing table. Each route table consists of a pile of k barrels, the so-called K bucket, is the bucket can only put K nodes, the default is 8. Buckets are saved in a way that is similar to a prefix tree. The equivalent of a maximum of 160-4 k barrels in a 8-bucket routing table.

3. According to the DHT protocol, each infohash has a position, so there is a distance between two Infohash, and two infohash distance can be expressed by XOR, that is, Infohash1 xor Infohash2, that is, the same as a high , they are close to each other, and vice versa, so that the distances of two nodes can be calculated quickly. What's the point of calculating this distance? In a DHT network, if the Infohash of a resource is closer to the infohash of a node, the more likely it is that the node has information about that resource. As you can imagine, because everyone uses the same distance algorithm to recursively ask a node that is close to the resource, and as long as the node responds, it gets a announce message, which means that the node with the resource Infohash has a greater probability of getting the resource Infohash

4. According to the above algorithm, the query in DHT is a jumping query, which can quickly cross the node bucket and close to the target node bucket. It is possible to jump very far in the distance, but only a small leap in the vicinity, because the more nodes in each node's routing table are saved, such as

5. In a DHT network when the crawler is not easy, not like ordinary reptiles, see resources can be actively climbed down, on the contrary, because the way to get resources (Get_peers, announce_peer) are passive, so the way the crawler has changed some, What a crawler does is to respond to queries from other nodes like a normal node, and to get the other nodes ' responses, and collect the data to get the job done. The only thing the crawler can do is try to get to know the other nodes as much as possible so that there are more other nodes to ask you.

6. Some people say, then I put the DHT crawler K bucket capacity K increase is not the opportunity to increase access to resources, in fact, before also analyzed, DHT crawler The most important source of information is all passive, because you can not increase the other people's K, so far from the node to save your own probability of the smaller, Of course, the probability of a distant node to request you is relatively small.

  • Contact Us

    The content source of this page is from Internet, which doesn't represent Alibaba Cloud's opinion; products and services mentioned on that page don't have any relationship with Alibaba Cloud. If the content of the page makes you feel confusing, please write us an email, we will handle the problem within 5 days after receiving your email.

    If you find any instances of plagiarism from the community, please send an email to: info-contact@alibabacloud.com and provide relevant evidence. A staff member will contact you within 5 working days.

    A Free Trial That Lets You Build Big!

    Start building with 50+ products and up to 12 months usage for Elastic Compute Service

    • Sales Support

      1 on 1 presale consultation

    • After-Sales Support

      24/7 Technical Support 6 Free Tickets per Quarter Faster Response

    • Alibaba Cloud offers highly flexible support services tailored to meet your exact needs.