標籤:
SequeceFile是Hadoop API提供的一種二進位檔案支援。這種二進位檔案直接將<key, value>對序列化到檔案中。可以使用這種檔案對小檔案合并,即將檔案名稱作為key,檔案內容作為value序列化到大檔案中。這種檔案格式有以下好處:
1). 支援壓縮,且可定製為基於Record或Block壓縮(Block級壓縮效能較優)
2). 本地化任務支援:因為檔案可以被切分,因此MapReduce任務時資料的本地化情況應該是非常好的。
3). 難度低:因為是Hadoop架構提供的API,商務邏輯側的修改比較簡單。
壞處:是需要一個合并檔案的過程,且合并後的檔案將不方便查看。
package test0820;import java.io.IOException;import java.io.InputStream;import java.net.URI;import org.apache.hadoop.conf.Configuration;import org.apache.hadoop.fs.FileStatus;import org.apache.hadoop.fs.FileSystem;import org.apache.hadoop.fs.Path;import org.apache.hadoop.io.IOUtils;import org.apache.hadoop.io.SequenceFile;import org.apache.hadoop.io.Text;public class TestSF { public static void main(String[] args) throws IOException, Exception{ Configuration conf = new Configuration(); FileSystem fs = FileSystem.get(new URI("hdfs://10.16.17.182:9000"), conf);
//輸入路徑:檔案夾
FileStatus[] files = fs.listStatus(new Path(args[0])); Text key = new Text(); Text value = new Text();
//輸出路徑:檔案
SequenceFile.Writer writer = SequenceFile.createWriter(fs, conf, new Path(args[1]),key.getClass() , value.getClass()); InputStream in = null; byte[] buffer = null; for(int i=0;i<files.length;i++){ key.set(files[i].getPath().getName()); in = fs.open(files[i].getPath()); buffer = new byte[(int) files[i].getLen()]; IOUtils.readFully(in, buffer, 0, buffer.length); value.set(buffer); IOUtils.closeStream(in); System.out.println(key.toString()+"\n"+value.toString()); writer.append(key, value); } IOUtils.closeStream(writer); }}
注意,待完善的地方:以Block方式壓縮。
MR案例:小檔案合并SequeceFile