CSV导入字符编码异常:ÅLAND ISLANDS显示为问号,求适配字符集
It’s frustrating when most special characters work but one stubbornly refuses to cooperate—let’s break down why this might be happening and how to fix it.
Possible Causes & Solutions
1. The File’s Actual Encoding Isn’t What You Think
Even if you’re specifying UTF-8, the CSV file might be saved in a different encoding that handles Å uniquely. Here’s how to check:
- Use a text editor like Notepad++: Open the file, look at the bottom-right corner—Notepad++ displays the current encoding. If Å shows correctly, use that encoding in your
InputStreamReader(e.g., if it says "UTF-8 with BOM", stick with UTF-8 but handle the BOM). - Check the raw bytes: Use a hex editor to look at the bytes for Å.
- UTF-8 Å should be
C3 85 - ISO-8859-1/Windows-1252 Å is
C5 - If you see
C5alone, reading with UTF-8 will turn it into a question mark since it’s an invalid UTF-8 sequence.
- UTF-8 Å should be
2. UTF-8 BOM Interference
Some files saved as UTF-8 include a Byte Order Mark (BOM: EF BB BF) at the start. Java’s standard UTF-8 decoder doesn’t automatically skip this, which can sometimes corrupt the first character. Fix this by using Apache Commons IO’s BOMInputStream to strip the BOM:
import org.apache.commons.io.input.BOMInputStream; // Wrap your FileInputStream in BOMInputStream InputStreamReader reader = new InputStreamReader( new BOMInputStream(new FileInputStream(file)), StandardCharsets.UTF_8 ); ICsvBeanReader beanReader = new CsvBeanReader(reader, ...);
3. The Issue Is in Display, Not Reading
You might be reading the character correctly, but your output method (console, database, UI) is using the wrong encoding:
- Console output: On Windows, the default console encoding isn’t UTF-8. Force it with:
In IDEs like IntelliJ/Eclipse, set the console encoding to UTF-8 in the settings.System.setOut(new PrintStream(System.out, true, StandardCharsets.UTF_8)); - Database storage: Ensure your database column uses a UTF-8-compatible charset (like
utf8mb4in MySQL) to avoid converting Å to a question mark when saving.
4. Mixed Encodings in the File
If the CSV was edited by multiple tools, parts of it might use different encodings. For example, Å could have been copied from an ISO-8859-1 source while other characters are UTF-8. The fix here is to re-save the entire file in UTF-8:
- Open the file in Notepad++, go to
Encoding > Convert to UTF-8, then save.
5. Auto-Detect the Encoding
If you’re unsure of the file’s encoding, use a library like juniversalchardet to detect it automatically:
import org.mozilla.universalchardet.UniversalDetector; import java.io.FileInputStream; import java.io.InputStreamReader; public String detectFileEncoding(File file) throws IOException { byte[] buf = new byte[4096]; FileInputStream fis = new FileInputStream(file); UniversalDetector detector = new UniversalDetector(null); int nread; while ((nread = fis.read(buf)) > 0 && !detector.isDone()) { detector.handleData(buf, 0, nread); } detector.dataEnd(); String encoding = detector.getDetectedCharset(); detector.reset(); fis.close(); return encoding; } // Then use it in your code: String detectedEncoding = detectFileEncoding(file); ICsvBeanReader beanReader = new CsvBeanReader( new InputStreamReader(new FileInputStream(file), detectedEncoding), new CsvPreference.Builder(CsvPreference.STANDARD_PREFERENCE).useQuoteMode(new AlwaysQuoteMode()).build() );
Final Tip
If all else fails, try opening the CSV in Google Sheets or Excel, re-save it as UTF-8 CSV, and then import it again. These tools often handle mixed encodings better than raw file readers.
内容的提问来源于stack exchange,提问作者acab05

