一、簡介
我們將一個正在啟動並執行程式稱為進程。每個進程都有它自己的系統狀態,包含記憶體狀態、開啟檔案清單、追蹤指令執行情況的程式指標以及一個儲存局部變數的調用棧。通常情況下,一個進程依照一個單序列控制流程順序執行,這個控制流程被稱為該進程的主線程。在任何給定的時刻,一個程式只做一件事情。
一個程式可以通過Python庫函數中的os或subprocess模組建立新進程(例如os.fork()或是subprocess.Popen())。然而,這些被稱為子進程的進程卻是獨立啟動並執行,它們有各自獨立的系統狀態以及主線程。因為進程之間是相互獨立的,因此它們同原有的進程並發執行。這是指原進程可以在建立子進程後去執行其它工作。
雖然進程之間是相互獨立的,但是它們能夠通過名為處理序間通訊(IPC)的機制進行相互連信。一個典型的模式是基於訊息傳遞,可以將其簡單地理解為一個純位元組的緩衝區,而send()或recv()操作原語可以通過諸如管道(pipe)或是網路通訊端(network socket)等I/O通道傳輸或接收訊息。還有一些IPC模式可以通過記憶體映射(memory-mapped)機制完成(例如mmap模組),通過記憶體映射,進程可以在記憶體中建立共用地區,而對這些地區的修改對所有的進程可見。
多進程能夠被用於需要同時執行多個任務的情境,由不同的進程負責任務的不同部分。然而,另一種將工作細分到任務的方法是使用線程。同進程類似,線程也有其自己的控制流程以及執行棧,但線程在建立它的進程之內運行,分享其父進程的所有資料和系統資源。當應用需要完成並發任務的時候線程是很有用的,但是潛在的問題是任務間必須分享大量的系統狀態。
當使用多進程或多線程時,作業系統負責調度。這是通過給每個進程(或線程)一個很小的時間片並且在所有活動任務之間快速迴圈切換來實現的,這個過程將CPU時間分割為小片段分給各個任務。例如,如果你的系統中有10個活躍的進程正在執行,作業系統將會適當的將十分之一的CPU時間分配給每個進程並且迴圈地在十個進程之間切換。當系統不止有一個CPU核時,作業系統能夠將進程調度到不同的CPU核上,保持系統負載平均以實現並存執行。
利用並發執行機制寫的程式需要考慮一些複雜的問題。複雜性的主要來源是關於同步和共用資料的問題。通常情況下,多個任務同時試圖更新同一個資料結構會造成髒資料和程式狀態不一致的問題(正式的說法是資源競爭的問題)。為瞭解決這個問題,需要使用互斥鎖或是其他相似的同步原語來標識並保護程式中的關鍵區段。舉個例子,如果多個不同的線程正在試圖同時向同一個檔案寫入資料,那麼你需要一個互斥鎖使這些寫操作依次執行,當一個線程在寫入時,其他線程必須等待直到當前線程釋放這個資源。
Python中的並發編程
Python長久以來一直支援不同方式的並發編程,包括線程、子進程以及其他利用產生器(generator function)的並發實現。
Python在大部分系統上同時支援訊息傳遞和基於線程的並發編程機制。雖然大部分程式員對線程介面更為熟悉,但是Python的線程機制卻有著諸多的限制。Python使用了內部全域解譯器鎖(GIL)來保證安全執行緒,GIL同時只允許一個線程執行。這使得Python程式就算在多核系統上也只能在單個處理器上運行。Python界關於GIL的爭論儘管很多,但在可預見的未來卻沒有將其移除的可能。
Python提供了一些很精巧的工具用於管理基於線程和進程的並行作業。即使是簡單地程式也能夠使用這些工具使得任務並發進行從而加快運行速度。subprocess模組為子進程的建立和通訊提供了API。這特別適合運行與文本相關的程式,因為這些API支援通過新進程的標準輸入輸出通道傳送資料。signal模組將UNIX系統的訊號量機制暴露給使用者,用以在進程之間傳遞事件資訊。訊號是非同步處理的,通常有訊號到來時會中斷程式當前的工作。訊號機制能夠實現粗粒度的訊息傳遞系統,但是有其他更可靠的進程內通訊技術能夠傳遞更複雜的訊息。threading模組為並行作業提供了一系列進階的,物件導向的API。Thread對象們在一個進程內並發地運行,分享記憶體資源。使用線程能夠更好地擴充I/O密集型的任務。multiprocessing模組同threading模組類似,不過它提供了對於進程的操作。每個進程類是真實的作業系統進程,並且沒有共用記憶體資源,但multiprocessing模組提供了進程間共用資料以及傳遞訊息的機制。通常情況下,將基於線程的程式改為基於進程的很簡單,只需要修改一些import聲明即可。
Threading模組樣本
以threading模組為例,思考這樣一個簡單的問題:如何使用分段並行的方式完成一個大數的累加。
import threading class SummingThread(threading.Thread): def __init__(self, low, high): super(SummingThread, self).__init__() self.low = low self.high = high self.total = 0 def run(self): for i in range(self.low, self.high): self.total += i thread1 = SummingThread(0, 500000)thread2 = SummingThread(500000, 1000000)thread1.start() # This actually causes the thread to runthread2.start()thread1.join() # This waits until the thread has completedthread2.join()# At this point, both threads have completedresult = thread1.total + thread2.totalprint(result)
自訂Threading類庫
我寫了一個便於使用threads的小型Python類庫,包含了一些有用的類和函數。
關鍵參數:
* do_threaded_work – 該函數將一系列給定的任務分配給對應的處理函數(分配順序不確定)
* ThreadedWorker – 該類建立一個線程,它將從一個同步的工作隊列中拉取工作任務並將處理結果寫入同步結果隊列
* start_logging_with_thread_info – 將線程id寫入所有日誌訊息。(依賴日誌環境)
* stop_logging_with_thread_info – 用於將線程id從所有的日誌訊息中移除。(依賴日誌環境)
import threadingimport logging def do_threaded_work(work_items, work_func, num_threads=None, per_sync_timeout=1, preserve_result_ordering=True): """ Executes work_func on each work_item. Note: Execution order is not preserved, but output ordering is (optionally). Parameters: - num_threads Default: len(work_items) --- Number of threads to use process items in work_items. - per_sync_timeout Default: 1 --- Each synchronized operation can optionally timeout. - preserve_result_ordering Default: True --- Reorders result_item to match original work_items ordering. Return: --- list of results from applying work_func to each work_item. Order is optionally preserved. Example: def process_url(url): # TODO: Do some work with the url return url urls_to_process = ["http://url1.com", "http://url2.com", "http://site1.com", "http://site2.com"] # process urls in parallel result_items = do_threaded_work(urls_to_process, process_url) # print(results) print(repr(result_items)) """ global wrapped_work_func if not num_threads: num_threads = len(work_items) work_queue = Queue.Queue() result_queue = Queue.Queue() index = 0 for work_item in work_items: if preserve_result_ordering: work_queue.put((index, work_item)) else: work_queue.put(work_item) index += 1 if preserve_result_ordering: wrapped_work_func = lambda work_item: (work_item[0], work_func(work_item[1])) start_logging_with_thread_info() #spawn a pool of threads, and pass them queue instance for _ in range(num_threads): if preserve_result_ordering: t = ThreadedWorker(work_queue, result_queue, work_func=wrapped_work_func, queue_timeout=per_sync_timeout) else: t = ThreadedWorker(work_queue, result_queue, work_func=work_func, queue_timeout=per_sync_timeout) t.setDaemon(True) t.start() work_queue.join() stop_logging_with_thread_info() logging.info('work_queue joined') result_items = [] while not result_queue.empty(): result = result_queue.get(timeout=per_sync_timeout) logging.info('found result[:500]: ' + repr(result)[:500]) if result: result_items.append(result) if preserve_result_ordering: result_items = [work_item for index, work_item in result_items] return result_items class ThreadedWorker(threading.Thread): """ Generic Threaded Worker Input to work_func: item from work_queue Example usage: import Queue urls_to_process = ["http://url1.com", "http://url2.com", "http://site1.com", "http://site2.com"] work_queue = Queue.Queue() result_queue = Queue.Queue() def process_url(url): # TODO: Do some work with the url return url def main(): # spawn a pool of threads, and pass them queue instance for i in range(3): t = ThreadedWorker(work_queue, result_queue, work_func=process_url) t.setDaemon(True) t.start() # populate queue with data for url in urls_to_process: work_queue.put(url) # wait on the queue until everything has been processed work_queue.join() # print results print repr(result_queue) main() """ def __init__(self, work_queue, result_queue, work_func, stop_when_work_queue_empty=True, queue_timeout=1): threading.Thread.__init__(self) self.work_queue = work_queue self.result_queue = result_queue self.work_func = work_func self.stop_when_work_queue_empty = stop_when_work_queue_empty self.queue_timeout = queue_timeout def should_continue_running(self): if self.stop_when_work_queue_empty: return not self.work_queue.empty() else: return True def run(self): while self.should_continue_running(): try: # grabs item from work_queue work_item = self.work_queue.get(timeout=self.queue_timeout) # works on item work_result = self.work_func(work_item) #place work_result into result_queue self.result_queue.put(work_result, timeout=self.queue_timeout) except Queue.Empty: logging.warning('ThreadedWorker Queue was empty or Queue.get() timed out') except Queue.Full: logging.warning('ThreadedWorker Queue was full or Queue.put() timed out') except: logging.exception('Error in ThreadedWorker') finally: #signals to work_queue that item is done self.work_queue.task_done() def start_logging_with_thread_info(): try: formatter = logging.Formatter('[thread %(thread)-3s] %(message)s') logging.getLogger().handlers[0].setFormatter(formatter) except: logging.exception('Failed to start logging with thread info') def stop_logging_with_thread_info(): try: formatter = logging.Formatter('%(message)s') logging.getLogger().handlers[0].setFormatter(formatter) except: logging.exception('Failed to stop logging with thread info')
使用樣本
from test import ThreadedWorkerfrom queue import Queue urls_to_process = ["http://facebook.com", "http://pypix.com"] work_queue = Queue()result_queue = Queue() def process_url(url): # TODO: Do some work with the url return url def main(): # spawn a pool of threads, and pass them queue instance for i in range(5): t = ThreadedWorker(work_queue, result_queue, work_func=process_url) t.setDaemon(True) t.start() # populate queue with data for url in urls_to_process: work_queue.put(url) # wait on the queue until everything has been processed work_queue.join() # print results print(repr(result_queue)) main()