Pythongreenlet implementation principles and examples

Source: Internet
Author: User
This article mainly introduces the implementation principles and examples of Pythongreenlet. greenlet is a parallel processing library in Python. if you need it, refer to the parallel development technology that has recently started to study Python, including multithreading, multi-process, coroutine, etc. I gradually sorted out some online materials, and today I sorted out greenlet-related materials.

Technical Background of concurrent processing

Parallel processing is currently very important, because in many cases, parallel computing can greatly improve the system throughput, especially in the era of multi-core multi-processor, so the ancient languages like lisp have been taken over again, and functional programming has become increasingly popular. This section describes a python library for Parallel Processing: greenlet. Python has a very famous library called stackless, which is used for concurrent processing. it mainly gets something called a tasklet microthread. The biggest difference between greenlet and stackless is that, is it lightweight? Not enough. The biggest difference is that greenlet requires you to handle thread switching. that is to say, you need to specify which greenlet to execute and which greenlet to execute.

Implementation mechanism of greenlet

I used to develop web programs using python and used the fastcgi mode all the time. then, multiple threads are started in each process for request processing. one problem here is that the response time of each request must be extremely short. Otherwise, the server will reject the service as long as multiple requests are slow, because no thread can respond to the request. we usually perform performance tests when our services are launched, so there is no major problem under normal conditions. however, it is impossible to test all scenarios. once it appears, the user will not respond for a long time. partially unavailable, resulting in all unavailability. later it was converted to coroutine, The greenlet in python. therefore, we have a simple understanding of its implementation mechanism.

Each greenlet is only a python object (PyGreenlet) in heap. Therefore, it is no problem for a process to create millions or even millions of greenlets.

The code is as follows:


Typedef struct _ greenlet {
PyObject_HEAD
Char * stack_start;
Char * stack_stop;
Char * stack_copy;
Intptr_t stack_saved;
Struct _ greenlet * stack_prev;
Struct _ greenlet * parent;
PyObject * run_info;
Struct _ frame * top_frame;
Int recursion_depth;
PyObject * weakreflist;
PyObject * exc_type;
PyObject * exc_value;
PyObject * exc_traceback;
PyObject * dict;
} PyGreenlet;

Each greenlet is actually a function and the context for saving the function execution. for a function, the context is its stack .. all greenlets of the same process share the user stack allocated by the same operating system. therefore, greenlet can only use the global stack at the same time without conflicting stack data. greenlet uses stack_stop and stack_start to store the bottom and top of its stack. if the stack_stop of The greenlet to be executed overlaps with the greenlet in the current stack, we need to temporarily save the data in these overlapping greenlet stacks to heap. the storage location is recorded by stack_copy and stack_saved, so that the stack stack_stop and stack_start are Copied back from heap during restoration. otherwise, the stack data will be damaged. therefore, the greenlet created by the application is implemented concurrently by constantly copying data to the heap or copying data from the heap to the stack. coroutine is very comfortable for io-type applications.

The following is a simple stack space model of greenlet (from greenlet. c)

The code is as follows:


A PyGreenlet is a range of C stack addresses that must be
Saved and restored in such a way that the full range of
Stack contains valid data when we switch to it.

Stack layout for a greenlet:

| ^ |
| Older data |
|
Stack_stop. | _______________ |
. |
. | Greenlet data |
. | In stack |
. * | _______________ |... _____________ Stack_copy + stack_saved
. |
. | Data | greenlet data |
. | Unrelated | saved |
. | To | in heap |
Stack_start. | this |... | _________ | stack_copy
| Greenlet |
|
| Newer data |
| Vvv |

The following is a simple greenlet code.

The code is as follows:


From greenlet import greenlet

Def test1 ():
Print 12
Gr2.switch ()
Print 34

Def test2 ():
Print 56
Gr1.switch ()
Print 78

Gr1 = greenlet (test1)
Gr2 = greenlet (test2)
Gr1.switch ()

The coroutine discussed currently is generally supported by programming languages. Currently, I know languages that support coroutine include python, lua, go, erlang, scala, and rust. The coroutine differs from the thread in that the coroutine is not switched by the operating system, but by the programmer code. that is to say, the switching is controlled by the programmer, so that there is no so-called thread security problem.

All coroutines share the context of the entire process, so that the exchange between coroutines is very convenient.

Compared with the second solution (I/O multiplexing), it makes the program written using coroutine more intuitive, rather than splitting a complete process into multiple managed events for processing. The disadvantage of coroutine may be that it cannot take advantage of multi-core advantages. However, this can be solved through coroutine + process.

Coroutine can be used to process concurrency to improve performance, or to implement a state machine to simplify programming. I use the second one. At the end of last year, I came into contact with python and learned about the concept of coroutine in python. later, I came into contact with yield through pycon china2011. greenlet is also a coroutine solution, and in my opinion it is a more available solution, it is especially used to process state machines.

At present, this part has been basically completed, and I will take some time to summarize it later.

Summary:

1) multi-process can take advantage of multi-core advantages, but inter-process communication is troublesome. In addition, the increase in the number of processes will degrade the performance and lead to high process switching costs. The complexity of the program process is lower than that of I/O multiplexing.

2) I/O multiplexing is used to process multiple logical flows within a process. process switching is not required, and the performance is high. In addition, information sharing between flows is simple. However, the advantages of multi-core cannot be used. In addition, the program process is cut into small pieces by event processing, which is complicated and difficult to understand.

3) threads run within a process and are scheduled by the operating system. switching costs are low. In addition, they share the virtual address space of the process and share information between threads. However, thread security issues lead to steep thread learning curves and are prone to errors.

4) coroutine is provided by programming languages and is controlled by programmers. Therefore, there is no thread security problem and can be used to process state machines and concurrent requests. However, multi-core advantages cannot be used.

The above four solutions can be used together. I am optimistic about the process + coroutine mode.

Contact Us

The content source of this page is from Internet, which doesn't represent Alibaba Cloud's opinion; products and services mentioned on that page don't have any relationship with Alibaba Cloud. If the content of the page makes you feel confusing, please write us an email, we will handle the problem within 5 days after receiving your email.

If you find any instances of plagiarism from the community, please send an email to: info-contact@alibabacloud.com and provide relevant evidence. A staff member will contact you within 5 working days.

A Free Trial That Lets You Build Big!

Start building with 50+ products and up to 12 months usage for Elastic Compute Service

  • Sales Support

    1 on 1 presale consultation

  • After-Sales Support

    24/7 Technical Support 6 Free Tickets per Quarter Faster Response

  • Alibaba Cloud offers highly flexible support services tailored to meet your exact needs.