Common strategies for "Go" distributed database split-table libraries

Source: Internet
Author: User

Reprint: http://www.cnblogs.com/VipBin/archive/2011/07/12/2104690.html

In large-capacity, high-load web systems, a series of split databases can effectively improve database capacity and performance. In the early stages of the beginner program, programmers usually like to design a library and a single function table structure according to the traditional database design pattern, which can cause serious performance problems and maintenance problems after the amount of data and concurrency reaches a certain level. It is very painful to start with the problem, so it should be considered when the system is set up.

At present, some database policies are based on a single library structure, and then synchronously distributed to several servers to achieve read and write separation. Individuals find such tactics very clumsy, or try to separate them well, otherwise the memory of each machine is easily overrun.

Generally only the data volume is larger than the split of the table, this should have no objection; there is a possible maintenance of the more important table, such as the article directory table, if there is the possibility to pour data from other systems, but also to remove the data when the table is accidentally damaged, found that forgot to back up, That is really want to cry without tears.

Here's an analysis:

One, TIME structure

If the business system is more time-sensitive, such as the news release system of the article table, the database can be designed into a temporal structure, by time divided into several structures:

1) Flat Plate type

Table similar to:
article_200901
article_200902
article_200903

In years or months can be self-determined, but with the date of the table is too much, there is no need. The general advice is to be on a monthly basis.

The difficulty with this kind of division is that if I want to list 20 data, and there are 2 of them in the three tables, then it is likely that the business will require reading three. If the time is long, there are dozens of tables, and each table is 0, that is not to read the entire system of the table can be? In addition to this structure, it is difficult to make paging.

Primary key: In this system, the primary key is a 13-bit time stamp with milliseconds, do not use automatic numbering, otherwise it is difficult to navigate through the primary key to the table, you can also take time on the query, but more cumbersome.

2) Archival type

Table similar to:
Article_old
Article_new

In order to solve the shortcomings of the flat plate, you can use the time-archive design, you can see that the system has only two tables. One is the old article table, one is a new article table, the new article table put 2 months of information, each day regularly 2 months in the first day of the article is classified into the old table. This can solve the performance problem, because the general news release system reads the new content, the old content read less, the second can be tactful to solve functional problems, such as the problem of flat-panel, in the filing of the most also only need to read 2 tables to complete.

The drawback of the archive is that the old table capacity is still relatively large, and if the business allows it, the older content in the old table can be re-archived or cleaned up directly.

Second, the structure of the forum

If according to the section of the article to split the table, such as news, sports section to split the table, on the one hand can make each table data volume separation, on the other hand, the interaction between the various sections can be minimized. If the data sheet in the news section is damaged or needs to be maintained, it will not affect the normal work of the sports section, thus reducing the risk. The structure of the forum is often used for such systems as BBS.

There are several methods of plate structure:

1) corresponding type

For the number of sections, and a more fixed form, the direct correspondence is good. such as the news section, you can separate the list of news lists, news articles, and so on.

News_category
News_article
Sports_category
Sports_article

You can see that each section corresponds to a set of identical table structures, with the benefit of a glance. In the functional, because there are still some gaps between the sections, so the need for joint query is not much, development than the way of time structure to be easy.

Primary key: Still to consider, in this system, the primary key is the section timestamp, simple timestamp or automatic numbering can also be used, the query should remember to bring the section for positioning table.

2) Hot and cold type

The disadvantage of the counterpart is that if the number of sections is large and uncertain, the number of tables to be separated is too much. For example: Baidu Bar, if you press an entry a table design, how many tables?

In this way.

Tieba_ car
Tieba_ aircraft
Tieba_ Rocket
Tieba__unite

The list of cars and rockets is a popular watch, defined as the new section is placed in the Unite table, until its more than 10,000 main stickers to open the corresponding table structure. Because in this system, the unpopular section is certainly much more than the hot section, these unpopular sections are usually only a few posts, for them to open the table is too wasteful, and the number of hot sections and visits, and more than the unpopular section more, very characteristic.

Unite table can also be expanded into a hashtable, using the MD5 encoding of the entry, can be divided into n tables, I forget, MD5 before a can be divided into 36 tables, two are 1296 tables, enough.

Tieba_unite_ab
Tieba_unite_ac
...

Third, hash structure

Hash structure is usually used for blogs and other user-based occasions, in the blog such a system has a few features, 1 is a very large number of users, 2 is the number of articles issued by each user is less, 3 is the user sent the article irregular, 4 is each user hair is not much, but the total is still very big. Based on these characteristics, with any of the above mentioned in any kind of sub-table is not appropriate, a no fixed time is not suitable for use, a lot of two users, but also is unpopular, so it is not appropriate to use the section (user) demolition.

Hash structure in the above mentioned, since per user is not good to split directly, then a group of users into a table.

Blog_aa
Blog_ab
Blog_ac
...

As said above, MD5 take the first two-bit hash can reach 1296 tables, if not enough, then add one more, the total can reach 46656 tables, not enough?

The number of tables is too many, to create these tables is also very troublesome, you can consider in the program before the database insert, more than one sentence to judge the existence of the table and create a statement, very practical, and not very expensive.

Primary key: Still to consider, in this system, the primary key is the user ID timestamp, simple timestamp or automatic numbering can also be used, but the query should remember to bring the user name used to locate the table.

IV. structure of total score

The above structure, according to each business system, can come up with a lot of estimates. But now the Internet business is becoming more and more complex, sometimes, a single split method can not achieve the requirements, need several split programs implemented together, multi-pronged, this time the logic will let people around Halo. I have developed a system that simply mixes the hash structure with the time structure and finds the logic quite complex.

Therefore, in addition to splitting the table, according to the most original single-Library single table, and then build a general table, is a very advantageous architecture. In this architecture, each time to the database will write twice times the data, read the main dependency on the performance of the split table, the general table for the implementation of the difficult to implement after the completion of the function and for daily scheduled backups, and the total table and the table is a complete backup of each other, any one of the broken table or data is not normal, You can read the correct data from the master table and restore it, and vice versa.

In the overall structure, the challenge is the performance and maintainability of the general table. My plan is that the general table can adopt some service software and architecture relative to ensure stability, such as Oracle, or LVS Pgpool PostgreSQL, focus on ensuring the stability of the data, relative, the sub-table with the lightweight MySQL, the focus is on speed. The ability to use different software and programs for the total score list is also a major feature of the overall structure.

Summarize:

How to optimize the system by splitting the table, the most basic is to according to business requirements and characteristics of analysis. This article is only to provide a few basic methods, the specific work to first brain to think, must not mess up, with the wrong workload to add 10 times times OH.

Common strategies for "Go" distributed database split-table libraries

Contact Us

The content source of this page is from Internet, which doesn't represent Alibaba Cloud's opinion; products and services mentioned on that page don't have any relationship with Alibaba Cloud. If the content of the page makes you feel confusing, please write us an email, we will handle the problem within 5 days after receiving your email.

If you find any instances of plagiarism from the community, please send an email to: info-contact@alibabacloud.com and provide relevant evidence. A staff member will contact you within 5 working days.

A Free Trial That Lets You Build Big!

Start building with 50+ products and up to 12 months usage for Elastic Compute Service

  • Sales Support

    1 on 1 presale consultation

  • After-Sales Support

    24/7 Technical Support 6 Free Tickets per Quarter Faster Response

  • Alibaba Cloud offers highly flexible support services tailored to meet your exact needs.