You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R tesseract提取PDF文本输出文件编码为ANSI而非UTF-8的问题求助

Fixing UTF-8 Encoding Issue with tesseract + writeLines on Windows

First, let’s cut straight to the solution that should resolve your encoding problem. The core issue isn’t with tesseract (it does return UTF-8 text as expected) but with how writeLines handles encoding on Windows systems.

Modified Code to Force UTF-8 Output

Update your writeLines call to explicitly specify UTF-8 encoding—this overrides Windows’ default system code page behavior:

eng <- tesseract(language = "eng", options = list(tessedit_pageseg_mode = 1))
text <- ocr("example.pdf", engine = eng)
# Explicitly set UTF-8 encoding when writing
writeLines(text, con = "example.txt", encoding = "UTF-8")

For even more explicit control, you can use a file connection to define encoding upfront:

con <- file("example.txt", encoding = "UTF-8")
writeLines(text, con)
close(con)

Why This Happens

On Windows, the default text file encoding is tied to your system’s local code page (what Notepad calls "ANSI"). When you use writeLines without specifying an encoding, R falls back to this system default instead of preserving the UTF-8 encoding from tesseract.

Your observation that Encoding(text) returns "unknown" is normal for ASCII-compatible strings in R: R only marks strings as UTF-8 if they contain non-ASCII characters. To confirm text is valid UTF-8, use this check:

# Install the utf8 package if you haven’t already
install.packages("utf8")
utf8::utf8_valid(text)

This will return TRUE if tesseract correctly output UTF-8 text.

Verifying the Fix

After running the modified code:

  • Open example.txt in Notepad: Go to File > Save As—the encoding dropdown should now show "UTF-8".
  • In R, read the file back to confirm:
    read_text <- readLines("example.txt", encoding = "UTF-8")
    Encoding(read_text)
    
    This will return "UTF-8" (or "unknown" if all characters are ASCII, but the file itself will still be UTF-8 encoded).

Additional Notes

Since you’re running the latest versions of R, tesseract, and RStudio, package bugs aren’t the culprit here. This is purely a Windows-specific default behavior that we’ve overridden by explicitly setting the encoding during the write operation.

内容的提问来源于stack exchange,提问作者armentieres

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 11:42:27