Hadoop控制輸出檔案命名

來源:互聯網
上載者:User

   在一般情況下,Hadoop 每一個 Reducer 產生一個輸出檔案,檔案以

  part-r-00000、part-r-00001 的方式進行命名。如果需要人為的控制輸出檔案的命

  名或者每一個 Reducer 需要寫出多個輸出檔案時,可以採用 MultipleOutputs 類來

  完成。MultipleOutputs 採用輸出記錄的索引值對(output Key 和 output Value)或者

  任一字元串來產生輸出檔案的名字,檔案一般以 name-r-nnnnn 的格式進行命名,

  其中 name 是程式設定的任意名字;nnnnn 表示分區號。

  MultipleOutputs 的使用方式 的使用方式: :: :

  想要使用 MultipeOutputs,需要完成以下四個步驟:

  1. 在 Reducer 中聲明 MultipleOutputs 的變數

  private MultipleOutputs

  2. 在 Reducer 的 setup 函數中進行 MultipleOutputs 的初始化

  protected void setup(Context context)throws IOException, InterruptedException {

  multipleOutputs = new MultipleOutputs

  }

  3. 在 reduce 函數中進行輸出控制

  protected void reduce(Text key, Iterable values, Context context)throws IOException,

  InterruptedException {

  for (Text value : values) {

  multipleOutputs.write(NullWritable.get(), value, key.toString());

  }

  }

  4. 在 cleanup 函數中關閉輸出 MultipleOutputs

  protected void cleanup(Context context)throws IOException, InterruptedException {

  multipleOutputs.close();

  }

  注意:multipleOutputs.write(key, value, baseOutputPath)方法的第三個函數表明了該輸出所在的目錄(相對於使用者指定的輸出目錄)。如果baseOutputPath不包含檔案分隔字元“/”,那麼輸出的檔案格式為baseOutputPath-r-nnnnn(name-r-nnnnn);如果包含檔案分隔字元“/”,例如baseOutputPath=“029070-99999/1901/part”,那麼輸出檔案則為

聯繫我們

該頁面正文內容均來源於網絡整理,並不代表阿里雲官方的觀點,該頁面所提到的產品和服務也與阿里云無關,如果該頁面內容對您造成了困擾,歡迎寫郵件給我們,收到郵件我們將在5個工作日內處理。

如果您發現本社區中有涉嫌抄襲的內容,歡迎發送郵件至: info-contact@alibabacloud.com 進行舉報並提供相關證據,工作人員會在 5 個工作天內聯絡您,一經查實,本站將立刻刪除涉嫌侵權內容。

A Free Trial That Lets You Build Big!

Start building with 50+ products and up to 12 months usage for Elastic Compute Service

  • Sales Support

    1 on 1 presale consultation

  • After-Sales Support

    24/7 Technical Support 6 Free Tickets per Quarter Faster Response

  • Alibaba Cloud offers highly flexible support services tailored to meet your exact needs.