標籤:char return trace print fileinput dia string font bsp
一、擷取檔案的編碼格式
當我們在使用檔案輸入輸出資料流時,經常會出現亂碼問題,這通常是由於編碼格式導致的。
以複製一份檔案為例:
我們用輸入資料流(FileInputStream)讀取檔案,然後用輸出資料流(FileOutPutStream)重新寫入到另一個檔案,
如果源檔案的編碼格式和我們重新寫入時的編碼格式不一致,那麼就可能出現亂碼問題。
因此,我們需要擷取源檔案的編碼格式,以便在重新寫入時使用相同的編碼格式。
下面介紹一個簡單的方式準確擷取檔案的編碼格式:
一般地,我們根據檔案的前三個位元組就可以判斷該檔案是什麼編碼格式,如下:
EF BB BF UTF-8
FE FF UTF-16/UCS-2, little endian
FF FE UTF-16/UCS-2, big endian
FF FE 00 00 UTF-32/UCS-4, little endian.
00 00 FE FF UTF-32/UCS-4, big-endian.
因此,我麼讀取檔案的前幾個位元組就可用於判斷檔案的編碼格式,代碼如下:
import java.io.BufferedInputStream;
import java.io.File;
import java.io.FileInputStream;
public class GetFileEncode {
public static void main(String[] args) {
String filePath = "D:\\javaTest\\test.txt";
File sourceFile = new File(filePath);
getFilecharset(sourceFile);
}
private static String getFilecharset(File sourceFile) {
String charset = "GBK";
byte[] first3Bytes = new byte[3];
try {
boolean checked = false;
BufferedInputStream bis = new BufferedInputStream(new FileInputStream(sourceFile));
bis.mark(0);
int read = bis.read(first3Bytes, 0, 3);
if (read == -1) {
return charset; // 檔案編碼為 ANSI
} else if (first3Bytes[0] == (byte) 0xFF
&& first3Bytes[1] == (byte) 0xFE) {
charset = "UTF-16LE"; // 檔案編碼為 Unicode
checked = true;
} else if (first3Bytes[0] == (byte) 0xFE
&& first3Bytes[1] == (byte) 0xFF) {
charset = "UTF-16BE"; // 檔案編碼為 Unicode big endian
checked = true;
} else if (first3Bytes[0] == (byte) 0xEF
&& first3Bytes[1] == (byte) 0xBB
&& first3Bytes[2] == (byte) 0xBF) {
charset = "UTF-8"; // 檔案編碼為 UTF-8
checked = true;
}
bis.reset();
if (!checked) {
int loc = 0;
while ((read = bis.read()) != -1) {
loc++;
if (read >= 0xF0)
break;
if (0x80 <= read && read <= 0xBF) // 單獨出現BF以下的,也算是GBK
break;
if (0xC0 <= read && read <= 0xDF) {
read = bis.read();
if (0x80 <= read && read <= 0xBF) // 雙位元組 (0xC0 - 0xDF)
// (0x80
// - 0xBF),也可能在GB編碼內
continue;
else
break;
} else if (0xE0 <= read && read <= 0xEF) {// 也有可能出錯,但是幾率較小
read = bis.read();
if (0x80 <= read && read <= 0xBF) {
read = bis.read();
if (0x80 <= read && read <= 0xBF) {
charset = "UTF-8";
break;
} else
break;
} else
break;
}
}
}
bis.close();
} catch (Exception e) {
e.printStackTrace();
}
System.out.print(charset);
return charset;
}
}
常用技術匯總