Due to the increase in business volume, the data on both sides occupied GB of memory during reconciliation. Considering the business growth volume, I plan to make some modifications to the method of performing reconciliation after all the original full-day data is read, change the join mode to a stream-like join mode, as shown in the following figure:
If the sequence of A's output stream is basically the same as that of B's output stream, A better hash join effect can be obtained, but for A few N generations (N consecutive times failed to match) make some compensation for unmatched data to complete all matching.
However, the order of the output stream in A is very different from that in B, which may cause A vast majority of data to fail to match. In this case, when compensation is made, the entire method degrades to join B Based on A left, and then join A Based on B left.