【Hadoop】資料序列化系統Avro

來源:互聯網
上載者:User

標籤:

  • Avro簡介
    • schema
  • 檔案組成
    • Header與Datablock聲明代碼
    • 測試代碼
  • 序列化與還原序列化
    • specific
    • generic
  • 參考資料

Avro簡介

Avro是由Doug Cutting(Hadoop之父)建立的資料序列化系統,旨在解決Writeable類型的不足:缺乏語言的可移植性。為了支援跨語言,Avro的schema與語言的模式無關。有關Avro的更多特性請參看官方文檔 1。

Avro檔案的讀寫是依據schema而進行的。通常情況下,Avro的schema是用JSON編寫,而資料部分則是二進位格式編碼,並採用壓縮演算法對資料進行壓縮,以便減少傳輸量。

schema

schema中資料欄位的類型包括兩種

  • 原生類型(primitive types): null, boolean, int, long, float, double, bytes, and string
  • 複雜類型(complex types): record, enum, array, map, union, and fixed

複雜類型比較常用的record。這裡用[2]中twitter.avro檔案為例,開啟檔案後,檔案頭如下:

Objavro.codecnullavro.schemaò{“type”:”record”,”name”:”twitter_schema”,”namespace”:”com.miguno.avro”,”fields”:[{“name”:”username”,”type”:”string”,”doc”:”Name of the user account on Twitter.com”},{“name”:”tweet”,”type”:”string”,”doc”:”The content of the user’s Twitter message”},{“name”:”timestamp”,”type”:”long”,”doc”:”Unix epoch time in milliseconds”}],”doc:”:”A basic schema for storing Twitter messages”}

將schema格式化之後

{    "type": "record",    "name": "twitter_schema",    "namespace": "com.miguno.avro",    "fields": [        {            "name": "username", "type": "string",            "doc": "Name of the user account on Twitter.com"        },        {            "name": "tweet", "type": "string",            "doc": "The content of the user‘s Twitter message"        },        {            "name": "timestamp", "type": "long",            "doc": "Unix epoch time in milliseconds"        }    ],    "doc:": "A basic schema fostoring Twitter messages"}

其中,name是該JSON串的名字,type是指明name的類型,doc是對該name更為詳細的說明。

檔案組成3中的圖對Avro檔案進行詳細地描述,一個檔案由header與多個data block組成。header主要由MetaDatas與16位sync marker組成,MetaDatas中的資訊包含codec與schema;codec是data block中的資料採用的壓縮方式,為null(不壓縮)或者是deflate。deflate演算法是gzip所採用的壓縮演算法,就我自己感覺而言壓縮比在6倍以上(具體還沒研究過)。

其實每個data block間都會間隔一個sync marker,具體參看4。sync marker是為了用於mapReduce階段時檔案分割與同步;此外Avro本身是為了mapReduce而設計的。

Header與Datablock聲明代碼
//org.apache.avro.file.DataFileStream.java  public static final class Header {    Schema schema;    Map<String,byte[]> meta = new HashMap<String,byte[]>();    private transient List<String> metaKeyList = new ArrayList<String>();    byte[] sync = new byte[DataFileConstants.SYNC_SIZE]; //byte[16]    private Header() {}  }  static class DataBlock {    private byte[] data;    private long numEntries;    private int blockSize;    private int offset = 0;    private boolean flushOnWrite = true;    private DataBlock(long numEntries, int blockSize) {      this.data = new byte[blockSize];      this.numEntries = numEntries;      this.blockSize = blockSize;    }
測試代碼
DataFileReader<Void> reader =                  new DataFileReader<Void>(new FsInput(new Path("twitter.avro"), new Configuration()),                                           new GenericDatumReader<Void>());//print schemaSystem.out.println(reader.getSchema().toString(true));//print meta List<String> metaKeyList = reader.getMetaKeys();System.out.println(metaKeyList.toString());System.out.println(reader.getMetaString("avro.codec"));System.out.println(reader.getMetaString("avro.schema"));//print blockountreader.getBlockCount();//print the data in data blockSystem.out.println(reader.next());

可以看到meta中存放的是avro.codec, avro.schema。

序列化與還原序列化

官網上給出了兩種序列化方式:specific與generic。

specific
// Serialize user1, user2 and user3 to diskDatumWriter<User> userDatumWriter = new SpecificDatumWriter<User>(User.class);DataFileWriter<User> dataFileWriter = new DataFileWriter<User>(userDatumWriter);dataFileWriter.create(user1.getSchema(), new File("users.avro"));dataFileWriter.append(user1);dataFileWriter.append(user2);dataFileWriter.append(user3);dataFileWriter.close();// Deserialize Users from diskDatumReader<User> userDatumReader = new SpecificDatumReader<User>(User.class);DataFileReader<User> dataFileReader = new DataFileReader<User>(file, userDatumReader);User user = null;while (dataFileReader.hasNext()) {// Reuse user object by passing it to next(). This saves us from// allocating and garbage collecting many objects for files with// many items.user = dataFileReader.next(user);System.out.println(user);}

specific的方式是根據所產生的User類,提取出schema來進行Avro的解析。

generic
GenericRecord user1 = new GenericData.Record(schema);user1.put("name", "Alyssa");user1.put("favorite_number", 256);// Leave favorite color nullGenericRecord user2 = new GenericData.Record(schema);user2.put("name", "Ben");user2.put("favorite_number", 7);user2.put("favorite_color", "red");// Serialize user1 and user2 to diskFile file = new File("users.avro");DatumWriter<GenericRecord> datumWriter = new GenericDatumWriter<GenericRecord>(schema);DataFileWriter<GenericRecord> dataFileWriter = new DataFileWriter<GenericRecord>(datumWriter);dataFileWriter.create(schema, file);dataFileWriter.append(user1);dataFileWriter.append(user2);dataFileWriter.close();

generic的方式是預先產生了一個schema,然後再根據其解析。因為Avro檔案會將schema寫在檔案頭,所以在平常做解析時,generic的方式更為常見。

avro-tools的jar包提供了對Avro檔案豐富的操作,包括對Avro檔案進行切割,以用於做測試資料。

Available tools:      compile  Generates Java code for the given schema.       concat  Concatenates avro files without re-compressing.   fragtojson  Renders a binary-encoded Avro datum as JSON.     fromjson  Reads JSON records and writes an Avro data file.     fromtext  Imports a text file into an avro data file.      getmeta  Prints out the metadata of an Avro data file.    getschema  Prints out schema of an Avro data file.          idl  Generates a JSON schema from an Avro IDL file       induce  Induce schema/protocol from Java class/interface via reflection.   jsontofrag  Renders a JSON-encoded Avro datum as binary.      recodec  Alters the codec of a data file.  rpcprotocol  Output the protocol of a RPC service   rpcreceive  Opens an RPC Server and listens for one message.      rpcsend  Sends a single RPC message.       tether  Run a tethered mapreduce job.       tojson  Dumps an Avro data file as JSON, one record per line.       totext  Converts an Avro data file to a text file.  trevni_meta  Dumps a Trevni file‘s metadata as JSON.trevni_random  Create a Trevni file filled with random instances of a schema.trevni_tojson  Dumps a Trevni file as JSON.
參考資料
  1. Apache Avro documentation. ?
  2. miguno, avro-cli-examples. ?
  3. xyw_Eliot, Avro簡介. ?
  4. guibin, AVRO檔案結構分析. ?

著作權聲明:本文為博主原創文章,未經博主允許不得轉載。

【Hadoop】資料序列化系統Avro

聯繫我們

該頁面正文內容均來源於網絡整理,並不代表阿里雲官方的觀點,該頁面所提到的產品和服務也與阿里云無關,如果該頁面內容對您造成了困擾,歡迎寫郵件給我們,收到郵件我們將在5個工作日內處理。

如果您發現本社區中有涉嫌抄襲的內容,歡迎發送郵件至: info-contact@alibabacloud.com 進行舉報並提供相關證據,工作人員會在 5 個工作天內聯絡您,一經查實,本站將立刻刪除涉嫌侵權內容。

A Free Trial That Lets You Build Big!

Start building with 50+ products and up to 12 months usage for Elastic Compute Service

  • Sales Support

    1 on 1 presale consultation

  • After-Sales Support

    24/7 Technical Support 6 Free Tickets per Quarter Faster Response

  • Alibaba Cloud offers highly flexible support services tailored to meet your exact needs.