Write Ahead Log (WAL) for HBase--overall architecture, threading model "Go"

Source: Internet
Author: User

Transferred from: http://www.cnblogs.com/ohuang/p/5807543.html

The problem solved

The Write Ahead log (WAL) of HBase provides a highly concurrent, persistent log save and replay mechanism. Each business data write operation (Put/delete) is billed to the Wal before execution.

If there is an HBase server outage, you can replay the operations that were not completed before executing from the Wal.

This paper mainly discusses the Wal mechanism of HBase, how to solve these problems from the threading model and the message mechanism level:

1. Because multiple HBase clients can initiate concurrent business data write requests to an HBase region server, Wal also supports concurrent multi-threaded log writes. --Ensure that log writes are thread-safe and highly concurrent.

2. For a single hbase client, its log order in the Wal should be in the same order as the client-initiated business data write request.

(for the above two-point requirements, it is easy to think, with a queue is done.) See the architecture diagram below. )

3. To ensure high reliability, the log is not only written to the file system's memory cache, but should also be forced to disk (that is, Wal's sync operation) as soon as possible. But sync is too frequent and performance gets worse. So:

(1) Sync should be executed asynchronously in multiple background threads

(2) Frequent multiple sync, can be combined into one time sync--appropriate relaxation requirements for reliability, improve performance.

Schema diagram--threading model, message mechanism

Here is the HBase Wal architecture diagram I drew. I added a lot of annotations to the diagram, so this picture should be self-explanatory:

Region Server RPC Service thread

These threads handle the business data write requests made by the HBase client through the RPC service invocation (actually, the Google Protobuf service call). In the example, "Region Server RPC Service Thread 1" Did 3 row append operations, and a forced brush disk sync operation.

The sync operation is to ensure that the previous append operation (including the business data involved) is reliably recorded in the log on disk, and HBase can perform relatively unreliable complex operations, such as writing Memstore. -This is the semantics of write ahead.

Visible from the schema diagram, the concurrent append operation simply adds the append request object to the queue.

The queue here is a Lmax disrutpor ringbuffer (as I've described in this article), which you can simply understand as a lock-free high concurrency queue.

The specific code for append is as follows:

For sync operations:

(1) Put a Syncfuture object in the queue, representing a sync operation request.

Each syncfuture has a self-increasing sequence id--This is globally unique and is created by the Lmax disrutpor queue. Later Syncfuture's sequence ID is higher.

(2) Call Syncfuture.get () to block the wait until the background thread (Syncrunner in the schema diagram) notifies syncfuture to exit the block, indicating that the Wal log has been saved on disk.

Wal log consumption thread

In the Wal mechanism, there is only one Wal-Log consumer thread that gets append and sync operations from the queue. Such a multi-producer, single-consumer model determines the globally unique order in which the Wal logs are written concurrently.

1. For the acquired append operation, directly call Hadoop Sequence File Writer to take this append operation (including metadata and row key, family, qualifier, timestamp, Value and other business data) to the file.

So the Wal log file uses the Hadoop sequence file format. Of course, it can also be replaced by other storage formats, such as Avro.

The Hadoop sequence file format is no longer described here, and its main features are:

(1) binary format. Row key, family, qualifier, timestamp, value, etc. hbase byte[] data are written sequentially to the file.

(2) In the sequence file, every few lines, a 16-byte magic number is inserted as the delimiter. This way, if the file is damaged, causing a row to be incomplete, you can skip the line by this magic number delimiter and Continue reading the next complete line.

(3) Support compression. can be compressed by row. You can also press block compression (to make multiple lines into a block)

2. For the acquired sync operation, the thread pool submitted to the background syncrunner (see schema diagram above) is executed asynchronously.

The above this.syncrunners is the Syncrunner thread pool. It can be seen that by calculating Syncrunnerindex, a simple round-robin commit algorithm is adopted.

    • In addition, the Wal log consumption thread will attempt to collect a batch of Syncfuture objects (that is, the sync operation) and commit to syncrunner one at a time.

So, in the above code, you can see that the incoming offer () method is this.syncfutures this syncfutures[] array, not a single Syncfuture object.

Collect a batch of re-commit, the performance is better. However, the more Syncfuture objects a single batch needs to accumulate, the more timely sync is, the longer it will cause the foreground region Server RPC service thread to block on Syncfuture.get ().

Therefore, there is a balance between throughput and timeliness. HBase is more prone to high throughput in order to support the writing of massive amounts of data, as reflected in the comments below. The specific number of Syncfuture constitute a batch, there is a certain strategy, no longer in this statement.

Syncrunner Threads

1. Get a Syncfuture (code in the red box) from the queue that was submitted by the Wal log consumer thread.

2. Call the file system API to perform the sync () operation (code in the Blue box)

    • Combine multiple frequent sync () operations to improve performance.

As mentioned above, the Wal-Log consuming thread submits multiple syncfuture at a time. For this, the Syncrunner thread will only implement the sync operation represented by the most recent syncfuture (that is, the one with the largest sequence ID). And ignoring the previous syncfuture.

This is the code in the Green box.

3. If sync () completes, or because the merge mentioned above ignores a certain syncfuture, then Releasesyncfuture () ==> object.notify () is called to notify the Syncfuture block to exit.

The region Server RPC service thread that was previously blocked on Syncfuture.get () can continue to execute down.

At this point, the entire Wal-write process is complete.

Summarize

I think it's easier to think of the order in which the queue is used to coordinate and ensure the log writes when it is written to a thread concurrently.

However, providing the sync () API ensures the reliability of log writes while avoiding frequent sync () operations that affect performance. -This is one of the highlights of the HBase Wal implementation.

Later I study Wal's checkpoint and read the Wal replay mechanism, and then share with you.

Write Ahead Log (WAL) for HBase--overall architecture, threading model "Go"

Contact Us

The content source of this page is from Internet, which doesn't represent Alibaba Cloud's opinion; products and services mentioned on that page don't have any relationship with Alibaba Cloud. If the content of the page makes you feel confusing, please write us an email, we will handle the problem within 5 days after receiving your email.

If you find any instances of plagiarism from the community, please send an email to: info-contact@alibabacloud.com and provide relevant evidence. A staff member will contact you within 5 working days.

A Free Trial That Lets You Build Big!

Start building with 50+ products and up to 12 months usage for Elastic Compute Service

  • Sales Support

    1 on 1 presale consultation

  • After-Sales Support

    24/7 Technical Support 6 Free Tickets per Quarter Faster Response

  • Alibaba Cloud offers highly flexible support services tailored to meet your exact needs.