標籤:
- Avro簡介
- 檔案組成
- Header與Datablock聲明代碼
- 測試代碼
- 序列化與還原序列化
- 參考資料
Avro簡介
Avro是由Doug Cutting(Hadoop之父)建立的資料序列化系統,旨在解決Writeable類型的不足:缺乏語言的可移植性。為了支援跨語言,Avro的schema與語言的模式無關。有關Avro的更多特性請參看官方文檔 1。
Avro檔案的讀寫是依據schema而進行的。通常情況下,Avro的schema是用JSON編寫,而資料部分則是二進位格式編碼,並採用壓縮演算法對資料進行壓縮,以便減少傳輸量。
schema
schema中資料欄位的類型包括兩種
- 原生類型(primitive types): null, boolean, int, long, float, double, bytes, and string
- 複雜類型(complex types): record, enum, array, map, union, and fixed
複雜類型比較常用的record。這裡用[2]中twitter.avro檔案為例,開啟檔案後,檔案頭如下:
Objavro.codecnullavro.schemaò{“type”:”record”,”name”:”twitter_schema”,”namespace”:”com.miguno.avro”,”fields”:[{“name”:”username”,”type”:”string”,”doc”:”Name of the user account on Twitter.com”},{“name”:”tweet”,”type”:”string”,”doc”:”The content of the user’s Twitter message”},{“name”:”timestamp”,”type”:”long”,”doc”:”Unix epoch time in milliseconds”}],”doc:”:”A basic schema for storing Twitter messages”}
將schema格式化之後
{ "type": "record", "name": "twitter_schema", "namespace": "com.miguno.avro", "fields": [ { "name": "username", "type": "string", "doc": "Name of the user account on Twitter.com" }, { "name": "tweet", "type": "string", "doc": "The content of the user‘s Twitter message" }, { "name": "timestamp", "type": "long", "doc": "Unix epoch time in milliseconds" } ], "doc:": "A basic schema fostoring Twitter messages"}
其中,name是該JSON串的名字,type是指明name的類型,doc是對該name更為詳細的說明。
檔案組成3中的圖對Avro檔案進行詳細地描述,一個檔案由header與多個data block組成。header主要由MetaDatas與16位sync marker組成,MetaDatas中的資訊包含codec與schema;codec是data block中的資料採用的壓縮方式,為null(不壓縮)或者是deflate。deflate演算法是gzip所採用的壓縮演算法,就我自己感覺而言壓縮比在6倍以上(具體還沒研究過)。
其實每個data block間都會間隔一個sync marker,具體參看4。sync marker是為了用於mapReduce階段時檔案分割與同步;此外Avro本身是為了mapReduce而設計的。
Header與Datablock聲明代碼
//org.apache.avro.file.DataFileStream.java public static final class Header { Schema schema; Map<String,byte[]> meta = new HashMap<String,byte[]>(); private transient List<String> metaKeyList = new ArrayList<String>(); byte[] sync = new byte[DataFileConstants.SYNC_SIZE]; //byte[16] private Header() {} } static class DataBlock { private byte[] data; private long numEntries; private int blockSize; private int offset = 0; private boolean flushOnWrite = true; private DataBlock(long numEntries, int blockSize) { this.data = new byte[blockSize]; this.numEntries = numEntries; this.blockSize = blockSize; }
測試代碼
DataFileReader<Void> reader = new DataFileReader<Void>(new FsInput(new Path("twitter.avro"), new Configuration()), new GenericDatumReader<Void>());//print schemaSystem.out.println(reader.getSchema().toString(true));//print meta List<String> metaKeyList = reader.getMetaKeys();System.out.println(metaKeyList.toString());System.out.println(reader.getMetaString("avro.codec"));System.out.println(reader.getMetaString("avro.schema"));//print blockountreader.getBlockCount();//print the data in data blockSystem.out.println(reader.next());
可以看到meta中存放的是avro.codec, avro.schema。
序列化與還原序列化
官網上給出了兩種序列化方式:specific與generic。
specific
// Serialize user1, user2 and user3 to diskDatumWriter<User> userDatumWriter = new SpecificDatumWriter<User>(User.class);DataFileWriter<User> dataFileWriter = new DataFileWriter<User>(userDatumWriter);dataFileWriter.create(user1.getSchema(), new File("users.avro"));dataFileWriter.append(user1);dataFileWriter.append(user2);dataFileWriter.append(user3);dataFileWriter.close();// Deserialize Users from diskDatumReader<User> userDatumReader = new SpecificDatumReader<User>(User.class);DataFileReader<User> dataFileReader = new DataFileReader<User>(file, userDatumReader);User user = null;while (dataFileReader.hasNext()) {// Reuse user object by passing it to next(). This saves us from// allocating and garbage collecting many objects for files with// many items.user = dataFileReader.next(user);System.out.println(user);}
specific的方式是根據所產生的User類,提取出schema來進行Avro的解析。
generic
GenericRecord user1 = new GenericData.Record(schema);user1.put("name", "Alyssa");user1.put("favorite_number", 256);// Leave favorite color nullGenericRecord user2 = new GenericData.Record(schema);user2.put("name", "Ben");user2.put("favorite_number", 7);user2.put("favorite_color", "red");// Serialize user1 and user2 to diskFile file = new File("users.avro");DatumWriter<GenericRecord> datumWriter = new GenericDatumWriter<GenericRecord>(schema);DataFileWriter<GenericRecord> dataFileWriter = new DataFileWriter<GenericRecord>(datumWriter);dataFileWriter.create(schema, file);dataFileWriter.append(user1);dataFileWriter.append(user2);dataFileWriter.close();
generic的方式是預先產生了一個schema,然後再根據其解析。因為Avro檔案會將schema寫在檔案頭,所以在平常做解析時,generic的方式更為常見。
avro-tools的jar包提供了對Avro檔案豐富的操作,包括對Avro檔案進行切割,以用於做測試資料。
Available tools: compile Generates Java code for the given schema. concat Concatenates avro files without re-compressing. fragtojson Renders a binary-encoded Avro datum as JSON. fromjson Reads JSON records and writes an Avro data file. fromtext Imports a text file into an avro data file. getmeta Prints out the metadata of an Avro data file. getschema Prints out schema of an Avro data file. idl Generates a JSON schema from an Avro IDL file induce Induce schema/protocol from Java class/interface via reflection. jsontofrag Renders a JSON-encoded Avro datum as binary. recodec Alters the codec of a data file. rpcprotocol Output the protocol of a RPC service rpcreceive Opens an RPC Server and listens for one message. rpcsend Sends a single RPC message. tether Run a tethered mapreduce job. tojson Dumps an Avro data file as JSON, one record per line. totext Converts an Avro data file to a text file. trevni_meta Dumps a Trevni file‘s metadata as JSON.trevni_random Create a Trevni file filled with random instances of a schema.trevni_tojson Dumps a Trevni file as JSON.
參考資料
- Apache Avro documentation. ?
- miguno, avro-cli-examples. ?
- xyw_Eliot, Avro簡介. ?
- guibin, AVRO檔案結構分析. ?
著作權聲明:本文為博主原創文章,未經博主允許不得轉載。
【Hadoop】資料序列化系統Avro