使用readtext读取UTF-8编码CSV文件时乱码问题求助
I’ve run into this exact frustration before on Windows systems stuck with non-UTF-8 locales, so let’s break down what’s going wrong and how to fix it.
Why This Happens
Your system locale is set to English_United Kingdom.1252, which means Windows defaults to this encoding for most file operations. When you save a UTF-8 CSV from Excel, it actually saves it as UTF-8 with BOM (Byte Order Mark) by default—something readtext doesn’t automatically handle when you only specify encoding = "UTF-8". The worse double-encoding you see when writing back is because write.csv interacts poorly with non-UTF-8 system locales, leading to double conversion of special characters.
Step-by-Step Solutions
1. Read the CSV with UTF-8-BOM Encoding
The simplest fix is to explicitly tell readtext to expect the BOM that Excel adds to its UTF-8 exports:
text_raw <- readtext::readtext("path/test_encoding.csv", encoding = "UTF-8-BOM", text_field = "c_text") text_raw
This should correctly parse characters like ü, ï, and á without turning them into garbled ü or ü strings.
2. Fallback: Read with read.csv First, Then Convert to readtext Object
If the first method doesn’t work (rare, but possible), read the CSV into a regular data frame first with proper encoding, then convert it to a readtext object:
# Read the CSV correctly into a data frame df <- read.csv("path/test_encoding.csv", fileEncoding = "UTF-8-BOM", stringsAsFactors = FALSE) # Convert to a readtext object text_raw <- readtext::readtext(df, text_field = "c_text")
3. Fix Writing to Avoid Double Encoding
The write.csv function is notoriously finicky with encodings on Windows. Instead, use either data.table::fwrite (recommended for speed and reliability) or write.table with explicit encoding settings:
Option A: Use data.table::fwrite
library(data.table) fwrite(text_raw, "path/output.csv", encoding = "UTF-8", row.names = FALSE)
Option B: Use write.table
write.table( text_raw, "path/output.csv", sep = ",", fileEncoding = "UTF-8", row.names = FALSE, quote = TRUE # Matches write.csv's default quoting behavior )
Both of these will skip the double-encoding issue you encountered with write.csv.
Key Takeaways
- Excel’s "UTF-8 CSV" is actually UTF-8 with BOM—always specify
UTF-8-BOMwhen reading these files in R on Windows. - Avoid
write.csvfor UTF-8 files on non-UTF-8 locales;fwriteorwrite.tableare far more reliable. - Even if you can’t change your system locale, explicitly setting encodings at every read/write step will bypass most encoding headaches.
内容的提问来源于stack exchange,提问作者Tim Welsh

