Long time not to come here to write things, a bit of today, come here to write something recently encountered. Some time ago, a telecommunications business users due to a core production library several times the recent outage restart, multi-personnel involved in the fruitless, sent me an email, probably meaning that now the problem has caused more serious consequences, hoping to help intervene in the analysis, diagnosis and solve the problem. By the people who were involved in the problem before, it is now the third time that the outage has been restarted, about 2 months. After the first reboot, because the OPS did not obtain the valuable information at that time, there was no conclusion; the second time other database related personnel were targeted to a bug that might have been the version (11.2.0.4,3), and given the solution, the reason they solved it, Because the information of the bug was found in the alert.log of the database, it is firmly believed that this is the root of the problem, after the implementation of the program, we have a solid heart. It is not expected that, after a long time, the same fault is still reproduced, so far, sent me an email. By communicating with the person concerned and getting the only information that was available when the problem occurred (not all of the information obtained), I just heard them say that the problem occurred when it was strange, the system suddenly hang on the look, and during the operation of the DB or OS level, no reaction, They also suspect OS or db-level anomalies, and even suspect a hardware problem ... , of course, their suspicions are justified. Through the DB logs provided by operations personnel, a strange problem was found that the database was not the result of an automatic outage due to a failure during the problem, but it was a bit surprising that the database was shut down, and it was a surprise that others were trying to deny it, which was expected, and no one wanted to admit it, besides, One of them happened at around 23, and they used the time to refute me: who would manipulate the database at this point in time? Think about it too, this is just a clue, the following is the log information:
Will cluster because some factors actively restart the database? Because the OPS personnel provided little information, they took the OS-level message and further confirmed my guess:
So what caused the cluster to proactively restart the database? Keep looking at the AWR report provided by OPS, navigate to the exception process and the corresponding SQL as follows:
Feedback users, the user side quickly locate the problem, after processing, so far nearly half a year, the failure did not happen again, all normal.
One of the serious database failure cases caused by performance problems