"The analysis report is provided to you by hindsight; statistical analysis is provided to you by foresight; and data mining is provided by insight )".
For example.
You saw Sun Wukong fighting with Erlang, and then wrote an analysis report, saying Sun Wukong had obvious flexibility advantages and Erlang was superior in strength, so he started to be equal. As a result, the two men ran to bamboo, sun Wukong won because Sun Wukong made full use of the advantages of bamboo. This is called an analysis report.
Sun Wukong is about to fight with Erlang. A gambler is asking you to make predictions. You made a statistical analysis and found that the two fought 4567 times, of which Sun Wukong won 3456 times. In addition, Sun Wukong wins 89%, while Erlang wins 71%. Sun Wukong wins the trend. Because you have assumed the relationship between this victory and history and made a hypothesis based on experience. This is called statistical analysis.
You did not do anything, and asked the computer to perform Association Analysis on its own. Four factors are automatically found: birth, education, experience, and Singleton. Sun Wukong wins. Through computer analysis, we found that children from poor backgrounds are generally harder to practice than those from the Royal Family. People with rich combat experience have more opportunities because they are good at using the environment, poor children may have a higher level of effort; single people are always better than non-single people in the same environment. Sun Wukong is no less famous than Erlang, and has a wealth of fighting experience and is single. Therefore, Sun Wukong won this fight. This is called data mining.
The difference between data mining and OLAP is that it has no assumptions and allows computers to find the relationship behind it, which may be what you think or unexpected. For example, the data mining results show that in the 0.2 billion combat records, the match between Sun and Yang is always the victory of sun, and Sun Wukong is the name of sun. Therefore, Wukong is the victory. In reality, for example, for OLAP analysis, we look for people who do not pay the money to the telecom operator in a timely manner. In general, we will analyze that people with low incomes often do not pay in time. Through analysis, it is found that 71% of the poor who do not pay in time. Data Mining is different. It analyzes the cause by itself. The reason may be that people outside the Fifth Ring Road do not pay the money in time. These conclusions are of great value to the promotion work. For example, market research outside the fifth ring road shows that more cooperation channels need to be established to facilitate payment. This is the value of data mining.
Data Mining is indeed good and can bring insights, but not necessarily bring insights to customers? Where is the source of insight?
The preceding examples show that the "surname" dimension indicators are different from those obtained by the "origin, education, and experience" dimension, the results obtained by analyzing the dimensions of "poor and rich" and "regional distribution" are completely different. But is it true that the more dimension indicators are created, the better?
I remember that we made five business analysis models some time ago. Each cube has many dimensions, and each dimension has many metrics, during the internal presentation, Hu Tao and I both put forward their opinions. Our opinion is that the analysis indicators must be based on the application. Every time an analysis indicator is set up, we must know what it can do, what kind of indicator combination can produce what kind of analysis results, instead of simply taking out a field in our database structure as an indicator. What we need to do now is "Try it out before we get down to the bottom". First, we can thoroughly understand our own things and then dig deeper.