This is a creation in
Article, where the information may have evolved or changed.
The implementation of a function in the Catena (Timing storage engine) is controversial, and it gets a metricsource from the map based on the name specified. This function is called at least once for each insert operation, and the function call is more frequent and spans multiple threads in a real-world scenario, so we have to consider synchronization.
The function obtains a pointer to Metricsource from Map[string]*metricsource, based on the specified name, and creates one and returns if it is not obtained. The key point to note is that we will only insert this map.
The simple implementation is as follows: (To save space, omit the function header and return, just paste the important part)
var source *memorySourcevarboolp.lock// lock the mutexdefer p.lock// unlock the mutex at the endif source, present = p.sources[name]; !present { // The source wasn't found, so we'll create it. source = &memorySource{ name: name, metrics: map[string]*memoryMetric{}, } // Insert the newly created *memorySource. p.sources[name] = source}
After testing, the implementation can reach approximately 1,400,000 insertions/sec (via the concurrent call, Gomaxprocs is set to 4). It looks fast, but in fact it is slower than a single co-process, because there is a lock contention between multiple processes.
Let's simplify the situation to illustrate the problem, assuming that the two processes are getting "a", "B", and "a" and "B" are already present in the map. At runtime, a process acquires a lock, takes a pointer, unlocks it, and resumes execution, at which point the other process is stuck in acquiring the lock. Waiting for a lock to release is time-consuming, and the more the process is, the worse the performance.
One way to make it faster is to remove the lock control and ensure that only one of the threads accesses the map. This method is simple, but not scalable. Let's look at another simple approach and ensure thread safety and scalability.
var source *memorysourcevar present Boolif source, present = P.sources[name];!present {//Added this line//the Sour Ce wasn ' t found, so we ' ll Create it. P.Lock. Lock ()// lock the mutex defer p.Lock. Unlock ()//Unlock at the endif source, present = P.sources[name]; !present {Source = &memorysource{name:name, metrics:map[string]*memorymetric{}, } // Insert the newly created *memorysource. P.sources[name] = source}// if present is true, then another goroutine have ALR Eady inserted//The element we want, and source is set to what we want.} Added this line//Note if the source is present, we avoid the lock completely!
The implementation can reach 5,500,000 inserts/second, 3.93 times times faster than the first version. There are 4 processes in the run test, the results are basically consistent with the expected value.
This implementation is OK because we did not delete or modify the operation. We can safely use the pointer address in the CPU cache, but be aware that we still need to lock it. If not, one of the threads may already be in the process of being inserted when creating the insert source, and they will be in a competitive state. In this version we only lock in very few cases, so the performance has improved a lot.
John POTOCNY recommends removing the defer because it delays the unlock time (to unlock when the entire function returns), and the following gives an "ultimate" version:
varSOURCE *memorysourcevarPresentBOOLifSOURCE, present = P.sources[name];!present {//The source wasn ' t found, so we ' ll create it.P.Lock. Lock ()//Lock the mutex ifSOURCE, present = P.sources[name];!present {Source = &memorysource{name:name, metrics : map[string]*memorymetric{},}//Insert the newly created *memorysource.P.sources[name] = source} p.Lock. Unlock ()//Unlock the Mutex}//Note that if the source is present, we avoid the lock completely!
9,800,000 Insert/sec! Changed 4 lines to 7 times!! There is wood!!!
Update : (the original author is very good and progressive)
Is the above implementation correct? No! With Go Data Race Detector We can easily discover the conditions of the condition and we cannot guarantee the integrity of the map while reading and writing.
The following gives the non-existent conditions, thread safety, should be considered "correct" version. With Rwmutex, read operations are not locked and write operations are kept in sync.
var source *memorysourcevar present bool p.lock . Rlock () if source, present = P.sources[name];!present { The source wasn ' t found, so we ' ll create it. P.lock . Runlock () P.lock . Lock () if source, present = P.sources[name];!present {Source = &memoryso urce{name:name, Metrics:map[string ]*memorymetric{},} //Insert the newly created *memorysource. P.sources[name] = source} p.lock . Unlock ()} else {p.lock . Runlock ()}
Tested, this version of the performance of its previous version of 93.8%, in order to ensure the correctness of the premise to reach this is very good. Perhaps we can think that there is no comparability between them, because the previous version was wrong.
Translate original
English original