You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用readtext读取UTF-8编码CSV文件时乱码问题求助

Fixing UTF-8 Encoding Issues with readtext on Windows

I’ve run into this exact frustration before on Windows systems stuck with non-UTF-8 locales, so let’s break down what’s going wrong and how to fix it.

Why This Happens

Your system locale is set to English_United Kingdom.1252, which means Windows defaults to this encoding for most file operations. When you save a UTF-8 CSV from Excel, it actually saves it as UTF-8 with BOM (Byte Order Mark) by default—something readtext doesn’t automatically handle when you only specify encoding = "UTF-8". The worse double-encoding you see when writing back is because write.csv interacts poorly with non-UTF-8 system locales, leading to double conversion of special characters.

Step-by-Step Solutions

1. Read the CSV with UTF-8-BOM Encoding

The simplest fix is to explicitly tell readtext to expect the BOM that Excel adds to its UTF-8 exports:

text_raw <- readtext::readtext("path/test_encoding.csv", encoding = "UTF-8-BOM", text_field = "c_text")
text_raw

This should correctly parse characters like ü, ï, and á without turning them into garbled ü or ü strings.

2. Fallback: Read with read.csv First, Then Convert to readtext Object

If the first method doesn’t work (rare, but possible), read the CSV into a regular data frame first with proper encoding, then convert it to a readtext object:

# Read the CSV correctly into a data frame
df <- read.csv("path/test_encoding.csv", fileEncoding = "UTF-8-BOM", stringsAsFactors = FALSE)

# Convert to a readtext object
text_raw <- readtext::readtext(df, text_field = "c_text")

3. Fix Writing to Avoid Double Encoding

The write.csv function is notoriously finicky with encodings on Windows. Instead, use either data.table::fwrite (recommended for speed and reliability) or write.table with explicit encoding settings:

Option A: Use data.table::fwrite

library(data.table)
fwrite(text_raw, "path/output.csv", encoding = "UTF-8", row.names = FALSE)

Option B: Use write.table

write.table(
  text_raw,
  "path/output.csv",
  sep = ",",
  fileEncoding = "UTF-8",
  row.names = FALSE,
  quote = TRUE  # Matches write.csv's default quoting behavior
)

Both of these will skip the double-encoding issue you encountered with write.csv.

Key Takeaways

  • Excel’s "UTF-8 CSV" is actually UTF-8 with BOM—always specify UTF-8-BOM when reading these files in R on Windows.
  • Avoid write.csv for UTF-8 files on non-UTF-8 locales; fwrite or write.table are far more reliable.
  • Even if you can’t change your system locale, explicitly setting encodings at every read/write step will bypass most encoding headaches.

内容的提问来源于stack exchange,提问作者Tim Welsh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 05:00:09