輸入的InputFormat—-SequenceFileInputFormat

來源:互聯網
上載者:User

繼承關係:SequenceFileInputFormat  extends FileInputFormat  implements InputFormat  。

 

SequenceFileInputFormat 代碼如下(其實很簡單):

  /**   * 覆蓋了FileInputFormat的這個方法,FileInputFormat通過這個方法得到的FileStatus[]   * 長度就是將要啟動並執行map的長度,每個FileStatus對應一個檔案   */  @Override  protected FileStatus[] listStatus(JobConf job) throws IOException {    FileStatus[] files = super.listStatus(job);    /*調用父類的listStatus方法,然後進行了自己的處理,將得到的FileStatus[]遍曆一遍,            遇到檔案夾時候,看其是否為MapFile,是的話去除其中也是SequenceFile的data檔案,否則將檔案夾過濾掉。     */    for (int i = 0; i < files.length; i++) {      FileStatus file = files[i];      if (file.isDir()) {     // it's a MapFile        Path dataFile = new Path(file.getPath(), MapFile.DATA_FILE_NAME);        FileSystem fs = file.getPath().getFileSystem(job);        // use the data file        files[i] = fs.getFileStatus(dataFile);      }    }    return files;  }

 

下面看看FileInputFormat的listStatus(JobConf job)方法:

 protected FileStatus[] listStatus(JobConf job) throws IOException {      //得到job配置中的所有輸入路徑,路徑中是以 ,號隔開的 。      Path[] dirs = getInputPaths(job);      if (dirs.length == 0) {          throw new IOException("No input paths specified in job");      }    // get tokens for all the required FileSystems..    TokenCache.obtainTokensForNamenodes(job.getCredentials(), dirs, job);        List<FileStatus> result = new ArrayList<FileStatus>();    List<IOException> errors = new ArrayList<IOException>();        // creates a MultiPathFilter with the hiddenFileFilter and the    // user provided one (if any).    //處理一個過路檔案的filter,可以將輸入檔案夾中的一些檔案過濾掉    List<PathFilter> filters = new ArrayList<PathFilter>();    filters.add(hiddenFileFilter);    PathFilter jobFilter = getInputPathFilter(job);    if (jobFilter != null) {      filters.add(jobFilter);    }    PathFilter inputFilter = new MultiPathFilter(filters);    //對於每個輸入檔案夾進行遍曆    for (Path p: dirs) {      FileSystem fs = p.getFileSystem(job);       //得到這個輸入檔案下下的所有檔案(夾)      FileStatus[] matches = fs.globStatus(p, inputFilter);      if (matches == null) {        errors.add(new IOException("Input path does not exist: " + p));      } else if (matches.length == 0) {        errors.add(new IOException("Input Pattern " + p + " matches 0 files"));      } else {        //遍曆輸入問價夾下的每個檔案(夾)        for (FileStatus globStat: matches) {          if (globStat.isDir()) {            //檔案夾的話,將該檔案夾下的所有檔案和檔案夾添加到結果中            //****注意此處沒有再往下層遍曆,而是將檔案和檔案夾都返回到結果中 。            for(FileStatus stat: fs.listStatus(globStat.getPath(),                inputFilter)) {              result.add(stat);            }                    } else {            //檔案的話直接添加到result中,其實沒有任何判斷該檔案是否是輸入需要的格式等等            result.add(globStat);          }        }      }    }    if (!errors.isEmpty()) {      throw new InvalidInputException(errors);    }    LOG.info("Total input paths to process : " + result.size());     return result.toArray(new FileStatus[result.size()]);  }

 

是以總結SequenceFileInputFormat中輸出檔案的規律(假設輸入檔案夾是/input ):

1、輸入檔案夾中的檔案 ,即滿足:/input/***檔案 。

2、輸入檔案夾中子檔案夾中的檔案 ,/input/***/***檔案 。

3、輸入檔案夾中的子檔案夾的子檔案夾中的data檔案 ,/input/***/***/data檔案 ,該主要是針對MapFile的  。

 

得到一個個檔案後,怎麼將檔案對應到InputSplit(有可能一個file映射1個InputSplit,也可能映射幾個InputSplit),代碼見下:

 

/** Splits files returned by {@link #listStatus(JobConf)} when   * they're too big.*/   @SuppressWarnings("deprecation")  public InputSplit[] getSplits(JobConf job, int numSplits)    throws IOException {    FileStatus[] files = listStatus(job);        // Save the number of input files in the job-conf    job.setLong(NUM_INPUT_FILES, files.length);    long totalSize = 0;                           // compute total size    for (FileStatus file: files) {                // check we have valid files      if (file.isDir()) {        throw new IOException("Not a file: "+ file.getPath());      }      totalSize += file.getLen();    }    long goalSize = totalSize / (numSplits == 0 ? 1 : numSplits);    long minSize = Math.max(job.getLong("mapred.min.split.size", 1),                            minSplitSize);    // generate splits    ArrayList<FileSplit> splits = new ArrayList<FileSplit>(numSplits);    NetworkTopology clusterMap = new NetworkTopology();    //對於每個檔案該分割成多少InputSplit的處理    for (FileStatus file: files) {      Path path = file.getPath();      FileSystem fs = path.getFileSystem(job);      long length = file.getLen();      BlockLocation[] blkLocations = fs.getFileBlockLocations(file, 0, length);      //允許切割檔案時候,即允許將一個file切割成多個InputSplit 。      if ((length != 0) && isSplitable(fs, path)) {         long blockSize = file.getBlockSize();        //分塊的大小,一般是以快為單位,即會選擇blockSize ,其實會將按照block來分塊,這樣比較合適        long splitSize = computeSplitSize(goalSize, minSize, blockSize);        long bytesRemaining = length;        //按照分割設定(每塊大小),將檔案從從offset=0 到offset=length 分割成 (length/splitSize+1) 個(其實也並不是這些個啦 ~ ~ ,最後一塊的大小可以是splitSize*SPLIT_SLOP)         while (((double) bytesRemaining)/splitSize > SPLIT_SLOP) {          String[] splitHosts = getSplitHosts(blkLocations,               length-bytesRemaining, splitSize, clusterMap);          splits.add(new FileSplit(path, length-bytesRemaining, splitSize,               splitHosts));          bytesRemaining -= splitSize;        }        //返回splits        if (bytesRemaining != 0) {          splits.add(new FileSplit(path, length-bytesRemaining, bytesRemaining,                      blkLocations[blkLocations.length-1].getHosts()));        }      } else if (length != 0) {        String[] splitHosts = getSplitHosts(blkLocations,0,length,clusterMap);        splits.add(new FileSplit(path, 0, length, splitHosts));      } else {         //Create empty hosts array for zero length files        splits.add(new FileSplit(path, 0, length, new String[0]));      }    }    LOG.debug("Total # of splits: " + splits.size());    return splits.toArray(new FileSplit[splits.size()]);  }

ok ,一切都完了 ,但是你看上面的代碼可能會產生一個疑問:SequenceFile是儲存的一個個的key-value值,這樣分割檔案的話,會不會破壞原有的資料結構,即 要是某一個key-value被分到了兩個FileSplit腫麼辦 ?

見文章:http://www.cnblogs.com/serendipity/articles/2112613.html

 

 

個人感覺mapreduce的這個工具InputFormat比較亂,往往不看原始碼,你永遠也無法知道到底哪些檔案被選上了  。

聯繫我們

該頁面正文內容均來源於網絡整理,並不代表阿里雲官方的觀點,該頁面所提到的產品和服務也與阿里云無關,如果該頁面內容對您造成了困擾,歡迎寫郵件給我們,收到郵件我們將在5個工作日內處理。

如果您發現本社區中有涉嫌抄襲的內容,歡迎發送郵件至: info-contact@alibabacloud.com 進行舉報並提供相關證據,工作人員會在 5 個工作天內聯絡您,一經查實,本站將立刻刪除涉嫌侵權內容。

A Free Trial That Lets You Build Big!

Start building with 50+ products and up to 12 months usage for Elastic Compute Service

  • Sales Support

    1 on 1 presale consultation

  • After-Sales Support

    24/7 Technical Support 6 Free Tickets per Quarter Faster Response

  • Alibaba Cloud offers highly flexible support services tailored to meet your exact needs.