R tesseract提取PDF文本输出文件编码为ANSI而非UTF-8的问题求助
First, let’s cut straight to the solution that should resolve your encoding problem. The core issue isn’t with tesseract (it does return UTF-8 text as expected) but with how writeLines handles encoding on Windows systems.
Modified Code to Force UTF-8 Output
Update your writeLines call to explicitly specify UTF-8 encoding—this overrides Windows’ default system code page behavior:
eng <- tesseract(language = "eng", options = list(tessedit_pageseg_mode = 1)) text <- ocr("example.pdf", engine = eng) # Explicitly set UTF-8 encoding when writing writeLines(text, con = "example.txt", encoding = "UTF-8")
For even more explicit control, you can use a file connection to define encoding upfront:
con <- file("example.txt", encoding = "UTF-8") writeLines(text, con) close(con)
Why This Happens
On Windows, the default text file encoding is tied to your system’s local code page (what Notepad calls "ANSI"). When you use writeLines without specifying an encoding, R falls back to this system default instead of preserving the UTF-8 encoding from tesseract.
Your observation that Encoding(text) returns "unknown" is normal for ASCII-compatible strings in R: R only marks strings as UTF-8 if they contain non-ASCII characters. To confirm text is valid UTF-8, use this check:
# Install the utf8 package if you haven’t already install.packages("utf8") utf8::utf8_valid(text)
This will return TRUE if tesseract correctly output UTF-8 text.
Verifying the Fix
After running the modified code:
- Open
example.txtin Notepad: Go to File > Save As—the encoding dropdown should now show "UTF-8". - In R, read the file back to confirm:
This will returnread_text <- readLines("example.txt", encoding = "UTF-8") Encoding(read_text)"UTF-8"(or"unknown"if all characters are ASCII, but the file itself will still be UTF-8 encoded).
Additional Notes
Since you’re running the latest versions of R, tesseract, and RStudio, package bugs aren’t the culprit here. This is purely a Windows-specific default behavior that we’ve overridden by explicitly setting the encoding during the write operation.
内容的提问来源于stack exchange,提问作者armentieres

