UBIFS檔案系統分析5 – 檔案讀寫

來源:互聯網
上載者:User

檔案資料管理

傳統檔案系統,如ext2通過ext2_inode的i_block成員管理檔案的資料,i_block的直接塊和間接塊組成了一棵樹,對某個邏輯地址的讀取,需要在這個樹上找到相應的物理塊指標。

UBIFS為每一片資料建立一個data node,這一片資料一般指UBIFS_BLOCK_SIZE。data node被插入到wandering tree上,通過ino+type+blockno組成的key在wandering tree上尋找blockno對應的data node,因此可以確保檔案相鄰資料區塊的data node索引資訊是聚集的,避免了尋找索引節點導致的檔案讀效能下降。

UBIFS的預讀 bulk-read

bulk-read類似於檔案系統的read-ahead。預讀是檔案系統的一個最佳化技術,在要求讀取的資料基礎上,多讀取一些。這是假定檔案的讀取經常是連續訪問的,因此系統嘗試在使用者真正請求資料前就把資料讀入記憶體中。

Read-ahead是linux VFS實現的,並不需要底層具體檔案系統的支援,read-ahead在傳統的塊裝置檔案系統上工作的很好,但是對於UBIFS並不很好,UBIFS工作於UBI之上,UBI又在MTD之上,而MTD又是同步的,本身並沒有實現請求隊列,這就意味著VFS阻塞在UBIFS的讀上面直到所有預讀塊讀完。相反block-device API是非同步,reader不需要等待read-ahead.

VFS的預讀是為hard drives設計的,但是raw flash裝置和hard disk在尋道上花費巨大時間是完全不同的,所以這個技術僅限於HDDs而對於flash來講沒有必要

因此,VFS read-ahead對UBIFS來說不是改善而是拖累,所以UBIFS disable VFS的read-ahead。但是UBIFS有自己的預讀實現,我們稱之為bulk-read,可以在檔案系統mount option啟用build-read

有一些flash裝置,一次讀取大片資料要比分多次讀取資料區塊很多,比如,OneNand可以read-while-load如果讀取多個page的話,所以有時一次讀大片資料可以帶來效能的改善,這就是bulk-read所有的。

如果UBIFS注意到一個檔案正在連續讀取(至少3個連續4KiB塊被讀取了),並且UBIFS看到後面的資料駐留在相同的LEB上,UBIFS開始發送大的資料讀取請求,以追求更高的單位讀取速率。

例如,假定使用者連續的讀取一個檔案。並且很幸運的檔案是連續的存放在flash media上。假定LEB25包含的data node都屬於這個檔案愛女,並且data node在邏輯上(檔案內位移)和物理上(邏輯塊內位移)都是連續的。假定使用者從LEB25的offset 0 開始讀取,在這種情況下UBIFS讀取整個LEB25,然後用讀到的資料重新整理檔案cache,那麼當使用者請求下一個data node時,它已經在file cache中了

很明顯的是,bulk-read在某些情況下可能會造成系統變慢,所以在使用bulk-read時一定要小心,尤其在一個已經高度片段化的系統上。

fs/ubifs/file.c分析

ubifs data node 磁碟結構
/**
 * struct ubifs_data_node - data node.
 * @ch: common header
 * @key: node key
 * @size: uncompressed data size in bytes
 * @compr_type: compression type (%UBIFS_COMPR_NONE, %UBIFS_COMPR_LZO, etc)
 * @padding: reserved for future, zeroes
 * @data: data
 *           
 * Note, do not forget to amend 'zero_data_node_unused()' function when
 * changing the padding fields.
 */     
struct ubifs_data_node {
    struct ubifs_ch ch;
    __u8 key[UBIFS_MAX_KEY_LEN];
    __le32 size;
    __le16 compr_type;
    __u8 padding[2]; /* Watch 'zero_data_node_unused()' if changing! */
    __u8 data[];
} __attribute__ ((packed));
@key 資料節點的key,由檔案的inode no + key type + block number組成
@size 未壓縮的資料大小,一般情況下,size是UBIFS_BLOCK_SIZE,最後一塊可能小於UBIFS_BLOCK_SIZE
@data 壓縮的資料
@ch->len - sizeof(struct ubifs_data_node): @data的大小可以通過common header內的len減去ubifs_data_node結構的尺寸得到

  57 static int read_block(struct inode *inode, void *addr, unsigned int block,
  58               struct ubifs_data_node *dn)
  59 {
  60     struct ubifs_info *c = inode->i_sb->s_fs_info;
  61     int err, len, out_len;
  62     union ubifs_key key;
  63     unsigned int dlen;
  64
  65     data_key_init(c, &key, inode->i_ino, block);
  66     err = ubifs_tnc_lookup(c, &key, dn);
  67     if (err) {
  68         if (err == -ENOENT)
  69             /* Not found, so it must be a hole */
  70             memset(addr, 0, UBIFS_BLOCK_SIZE);
  71         return err;
  72     }
  73
  74     ubifs_assert(le64_to_cpu(dn->ch.sqnum) >
  75              ubifs_inode(inode)->creat_sqnum);
  76     len = le32_to_cpu(dn->size);
  77     if (len <= 0 || len > UBIFS_BLOCK_SIZE)
  78         goto dump;
  79    
  80     dlen = le32_to_cpu(dn->ch.len) - UBIFS_DATA_NODE_SZ;
  81     out_len = UBIFS_BLOCK_SIZE;
  82     err = ubifs_decompress(&dn->data, dlen, addr, &out_len,
  83                    le16_to_cpu(dn->compr_type));
  84     if (err || len != out_len)
  85         goto dump;
  86
  87     /*
  88      * Data length can be less than a full block, even for blocks that are
  89      * not the last in the file (e.g., as a result of making a hole and
  90      * appending data). Ensure that the remainder is zeroed out.
  91      */         
  92     if (len < UBIFS_BLOCK_SIZE)
  93         memset(addr + len, 0, UBIFS_BLOCK_SIZE - len);
  94
  95     return 0;
  96          
  97 dump:
  98     ubifs_err("bad data node (block %u, inode %lu)",
  99           block, inode->i_ino);
 100     dbg_dump_node(c, dn);
 101     return -EINVAL;
 102 }
read_block是readpage主要實現,@inode要讀取的檔案inode, @addr資料儲存的buffer, @block指明了要放問的檔案的塊號, dn輔助資料
65 首先用ino和blockno產生wandering tree的key
66 從樹中讀取data node
82 ubifs_decompress把dn->data的資料解壓到@addr中
92 ~ 93 如果解壓後的資料小於UBIFS_BLOCK_SIZE那麼把其他部分填充0, 注釋中描述了兩種情況導致解壓後的資料size小於UBIFS_BLOCK_SIZE
1. 檔案末尾
2. 檔案中間的hole

 700 /**
 701  * ubifs_do_bulk_read - do bulk-read.
 702  * @c: UBIFS file-system description object
 703  * @bu: bulk-read information
 704  * @page1: first page to read
 705  *
 706  * This function returns %1 if the bulk-read is done, otherwise %0 is returned.
 707  */
 708 static int ubifs_do_bulk_read(struct ubifs_info *c, struct bu_info *bu,
 709                   struct page *page1)
 710 {
 711     pgoff_t offset = page1->index, end_index;
 712     struct address_space *mapping = page1->mapping;
 713     struct inode *inode = mapping->host;
 714     struct ubifs_inode *ui = ubifs_inode(inode);
 715     int err, page_idx, page_cnt, ret = 0, n = 0;
 716     int allocate = bu->buf ? 0 : 1;
 717     loff_t isize;
 718
 719     err = ubifs_tnc_get_bu_keys(c, bu);
 720     if (err)
 721         goto out_warn;
 722
 723     if (bu->eof) {
 724         /* Turn off bulk-read at the end of the file */
 725         ui->read_in_a_row = 1;
 726         ui->bulk_read = 0;
 727     }
 728
 729     page_cnt = bu->blk_cnt >> UBIFS_BLOCKS_PER_PAGE_SHIFT;
 730     if (!page_cnt) {
 731         /*
 732          * This happens when there are multiple blocks per page and the
 733          * blocks for the first page we are looking for, are not
 734          * together. If all the pages were like this, bulk-read would
 735          * reduce performance, so we turn it off for a while.
 736          */
 737         goto out_bu_off;
 738     }
 739
 740     if (bu->cnt) {
 741         if (allocate) {
 742             /*
 743              * Allocate bulk-read buffer depending on how many data
 744              * nodes we are going to read.
 745              */
 746             bu->buf_len = bu->zbranch[bu->cnt - 1].offs +
 747                       bu->zbranch[bu->cnt - 1].len -
 748                       bu->zbranch[0].offs;
 749             ubifs_assert(bu->buf_len > 0);
 750             ubifs_assert(bu->buf_len <= c->leb_size);
 751             bu->buf = kmalloc(bu->buf_len, GFP_NOFS | __GFP_NOWARN);
 752             if (!bu->buf)
 753                 goto out_bu_off;
 754         }
 755
 756         err = ubifs_tnc_bulk_read(c, bu);
 757         if (err)
 758             goto out_warn;
 759     }
 760
 761     err = populate_page(c, page1, bu, &n);
 762     if (err)
 763         goto out_warn;
 764
 765     unlock_page(page1);
 766     ret = 1;
 767
 768     isize = i_size_read(inode);
 769     if (isize == 0)
 770         goto out_free;
 771     end_index = ((isize - 1) >> PAGE_CACHE_SHIFT);
 772
 773     for (page_idx = 1; page_idx < page_cnt; page_idx++) {
 774         pgoff_t page_offset = offset + page_idx;
 775         struct page *page;
 776
 777         if (page_offset > end_index)
 778             break;
 779         page = find_or_create_page(mapping, page_offset,
 780                        GFP_NOFS | __GFP_COLD);
 781         if (!page)
 782             break;
 783         if (!PageUptodate(page))
 784             err = populate_page(c, page, bu, &n);
 785         unlock_page(page);
 786         page_cache_release(page);
 787         if (err)
 788             break;
 789     }
 790
 791     ui->last_page_read = offset + page_idx - 1;
 792
 793 out_free:
 794     if (allocate)
 795         kfree(bu->buf);
 796     return ret;
 797
 798 out_warn:
 799     ubifs_warn("ignoring error %d and skipping bulk-read", err);
 800     goto out_free;
 801
 802 out_bu_off:
 803     ui->read_in_a_row = ui->bulk_read = 0;
 804     goto out_free;
 805 }
ubifs_do_bulk_read不僅讀取資料到@page1中,而且還會分配新的page並且把bulk-read讀取的資料儲存到這些新的page中
719 擷取這些data node的zbranch
741 ~ 754 需要分配bu->buf
756 ubifs_tnc_bulk_read 讀取@bu指定的flash資料到@bu->buf中
761 populate_page會把bu->buf內的資料解壓然後儲存到page中
773 ~ 789 為其他預讀的資料產生page cache,並存入讀取的資料

 807 /**
 808  * ubifs_bulk_read - determine whether to bulk-read and, if so, do it.
 809  * @page: page from which to start bulk-read.
 810  *
 811  * Some flash media are capable of reading sequentially at faster rates. UBIFS
 812  * bulk-read facility is designed to take advantage of that, by reading in one
 813  * go consecutive data nodes that are also located consecutively in the same
 814  * LEB. This function returns %1 if a bulk-read is done and %0 otherwise.
 815  */
 816 static int ubifs_bulk_read(struct page *page)
 817 {
 818     struct inode *inode = page->mapping->host;
 819     struct ubifs_info *c = inode->i_sb->s_fs_info;
 820     struct ubifs_inode *ui = ubifs_inode(inode);
 821     pgoff_t index = page->index, last_page_read = ui->last_page_read;
 822     struct bu_info *bu;
 823     int err = 0, allocated = 0;
 824
 825     ui->last_page_read = index;
 826     if (!c->bulk_read)
 827         return 0;
 828
 829     /*
 830      * Bulk-read is protected by @ui->ui_mutex, but it is an optimization,
 831      * so don't bother if we cannot lock the mutex.
 832      */
 833     if (!mutex_trylock(&ui->ui_mutex))
 834         return 0;
 835
 836     if (index != last_page_read + 1) {
 837         /* Turn off bulk-read if we stop reading sequentially */
 838         ui->read_in_a_row = 1;
 839         if (ui->bulk_read)
 840             ui->bulk_read = 0;
 841         goto out_unlock;
 842     }
 843
 844     if (!ui->bulk_read) {
 845         ui->read_in_a_row += 1;
 846         if (ui->read_in_a_row < 3)
 847             goto out_unlock;
 848         /* Three reads in a row, so switch on bulk-read */
 849         ui->bulk_read = 1;
 850     }
 851
 852     /*
 853      * If possible, try to use pre-allocated bulk-read information, which
 854      * is protected by @c->bu_mutex.
 855      */
 856     if (mutex_trylock(&c->bu_mutex))
 857         bu = &c->bu;
 858     else {
 859         bu = kmalloc(sizeof(struct bu_info), GFP_NOFS | __GFP_NOWARN);
 860         if (!bu)
 861             goto out_unlock;
 862
 863         bu->buf = NULL;
 864         allocated = 1;
 865     }
 866
 867     bu->buf_len = c->max_bu_buf_len;
 868     data_key_init(c, &bu->key, inode->i_ino,
 869               page->index << UBIFS_BLOCKS_PER_PAGE_SHIFT);
 870     err = ubifs_do_bulk_read(c, bu, page);
 871
 872     if (!allocated)
 873         mutex_unlock(&c->bu_mutex);
 874     else
 875         kfree(bu);
 876
 877 out_unlock:
 878     mutex_unlock(&ui->ui_mutex);
 879     return err;
 880 }

825 系統不支援bulk-read,直接返回
835 ~ 842 如果不是連續的讀取,那麼終止當前inode的bulk-read
844 ~ 850 連續的讀取了三個塊,那麼啟動bulk-read
867 預讀32 data node
868 ~ 870 初始化bu的key:檔案的ino + page的起始地址

 882 static int ubifs_readpage(struct file *file, struct page *page)
 883 {
 884     if (ubifs_bulk_read(page))
 885         return 0;
 886     do_readpage(page);
 887     unlock_page(page);
 888     return 0;
 889 }
ubifs_readpage是ubifs address_space_operations的readpage介面實現
884 首先通過ubifs特有的ubifs_bulk_read中去讀取給定的page,如果支援bulk read,則調用bulk read讀取
886 do_readpage讀取資料

聯繫我們

該頁面正文內容均來源於網絡整理,並不代表阿里雲官方的觀點,該頁面所提到的產品和服務也與阿里云無關,如果該頁面內容對您造成了困擾,歡迎寫郵件給我們,收到郵件我們將在5個工作日內處理。

如果您發現本社區中有涉嫌抄襲的內容,歡迎發送郵件至: info-contact@alibabacloud.com 進行舉報並提供相關證據,工作人員會在 5 個工作天內聯絡您,一經查實,本站將立刻刪除涉嫌侵權內容。

A Free Trial That Lets You Build Big!

Start building with 50+ products and up to 12 months usage for Elastic Compute Service

  • Sales Support

    1 on 1 presale consultation

  • After-Sales Support

    24/7 Technical Support 6 Free Tickets per Quarter Faster Response

  • Alibaba Cloud offers highly flexible support services tailored to meet your exact needs.