This is a creation in
Article, where the information may have evolved or changed.
About Cache
For backend developers, caching is a must-have skill. This is a panacea that you don't need to spend too much energy on to dramatically improve service performance. The premise is that you need to know how to use it so that you can maximize its effectiveness and suppress its side effects. This article describes the most appropriate way to add and update a cache.
Before you begin
This section describes what we have to do before we start adding caches. This step is very important, if not, it is possible to add a cache instead of adding.
Why use caching? For a service its performance bottlenecks are often in db, especially for traditional relational storage. When we create a table, we do not create an index of all the fields, which means that if we need to read the non-cached data we will take the data from the disk. This process takes at least more than 10 milliseconds. The cache is often memory-based, which is two orders of magnitude faster than DB read data. This is the root cause of the cache we use.
Then simply throw all the data into memory. No way. Memory this thing is fast, and it's expensive. It's a bit too wasteful to throw memory at every sell his grove g of data. According to the 28 law, we just need to find the most tight 20% on the line. This is very important. Otherwise you will have a worse cache effect.
There is a metric for caching, called Cache hit ratio. The high indicator indicates that most of the data we request comes from the cache. The higher the benefit of proving that we are adding caches.
Add cache
If you use some ORM tools in your usual situation, it is likely that you will not encounter these problems directly, but these questions need to be clearly understood before you add the cache. Some sort of generic routine. Let's take a look at:
Cache penetration
Cache penetration means accessing data that is not in a cache, but it does not exist in this data database. The general idea that we do not get data from the database does not trigger the cache operation. At this point, if a malicious attack, a large number of access to the cache directly through the database, to the backend services and databases to make great pressure or even downtime.
Solution:
-
Caches an empty object. If the cache is missing and there is no such object in the database, you can cache an empty object to the cache. If you use Redis, this key needs to be set up for a short time to prevent memory wastage.
-
Cache predictions. Predict if key exists. If the amount of cache can be used to determine the size of a hash, if the amount of large can be used to judge the filter.
Cache concurrency
Cache concurrency This scenario is easy to explain: multiple clients accessing a data that is not in the cache at the same time, each client executes the data set from the DB to the cache, resulting in cache concurrency.
Solution:
-
Cache warm-up. Add all the expected hot data to the cache in advance. Locating hot data is a complex matter that needs to be evaluated according to your own service access situation. This scheme can only reduce the number of concurrent cache occurrences can not be fully resisted.
-
Cache lock. If more than one client accesses a nonexistent cache, locks the logic before executing the load data and set cache, allowing only one client to execute the logic.
Cache anti-Avalanche
A cache avalanche is a service that is temporarily unavailable by the caching service, causing all requests to have direct access to the DB.
Solution:
-
Build a highly available cache system. The current common cache system Redis and memcache support a highly available deployment approach, so it is not a good time to consider whether or not to deploy in a highly available cluster mode.
-
Current limit. Netflix's Hystrix is a great tool to use when caching.
Update cache
In this section we will introduce the cache update policy. This part of the content is mainly from the Coolshell left ear mouse teacher, the text at the end of the original address, we can go to read.
Cache aside Pattern
Cache-aside-pattern.png
The idea is to update the database before the cache expires after the update succeeds. Another way is to fail the cache first, and then update the database. Let's compare the differences between the two ways.
First, look at the latter one. Imagine a scenario in which a client initiates an update operation when a cache execution fails. A read operation comes in and finds that the cache has no data and then takes the data from the database and puts it into the cache. The update operation continues to update the database. Dirty data is already cached in the cache.
So the first kind of problem arises? Theoretically, look at this: a client initiates an update operation, the B client initiates a read operation, and the cache is invalidated, and then it loads the data from the database (old data). A's update operation completes the failed cache, at which point the client reads the old data set to the cache. This is the case when dirty data is present, but the probability is very small.
Read/write Through Pattern
Read-write-through.png
Read Through: When reading data, if there is no data in the current cache, the usual operation is that the application goes to DB to load the data and then adds it to the cache. The difference between Read through is that we don't need to load the data in the application itself, and the cache layer will do something about it.
Write Through: When updating the data, if the cache is hit, the cache is updated and then cached to update the data to the database, and the database is updated directly if there is no hit cache.
This way the cache layer directly masks the DB, and the application only needs to deal with the cache. The advantage is that the application logic is simple and more efficient; The disadvantage is that the implementation of the cache layer is relatively complex.
Write back Pattern
Write-back.png
This is the most difficult way to achieve in three ways, it requires a dedicated storage to save the cache is dirty data, and read and write the cache when the dirty data synchronization. This can be used in scenarios where the data consistency requirement is not too high.
First, let's take a look at the read cache operation. If the cache hit is returned directly. If the cache is not hit, first to retrieve the key in the Strore is dirty, if not load the data, if the data should be flush to the storage, and then load the data. Next, Mark this key as not dirty, and return the result.
The process of writing data. If the hit cache updates the data and marks the record as dirty. If there is no hit, then go to the store to retrieve whether this can be dirty, if not the load data from the storage, update this data, if it is the current data flush to the storage, and then load data Update, and mark this record as dirty.
Reference