Python实现docx转纯文本代码触发UnicodeDecodeError问题求助
Hey there, let's break down what's going wrong with your code and how to fix that frustrating UnicodeDecodeError.
First, let's pinpoint the root issue in your original code:
You’re using
open(file)to read the DOCX file directly in your loop, but this is completely unnecessary—docx.Document(filename)already handles loading the DOCX file properly behind the scenes. When you useopen(file)without specifying an encoding, Python defaults to your system’s native charmap encoding (often limited on Windows), which can’t parse certain special characters in your DOCX file, triggering the decode error.
Here's the cleaned-up, error-free version of your code, with key improvements explained:
import docx import glob def getText(filename): doc = docx.Document(filename) fullText = [] for para in doc.paragraphs: fullText.append(para.text) return '\n'.join(fullText) # Loop through all DOCX files in the current folder for file in glob.glob('*.docx'): out_filename = file.replace('.docx', '-out.txt') # Fetch the processed text content directly from your function doc_content = getText(file) # Use a with-statement to handle file writing (auto-closes the file) with open(out_filename, 'w', encoding='utf8') as outfile: outfile.write(doc_content)
Let's walk through the critical fixes:
- Removed the redundant
for line in open(file):loop—this was the main cause of your encoding error, since we never need to read a DOCX file as raw plain text. - Switched to a
withstatement for writing the output file: this is Python best practice, as it automatically closes the file after writing, avoiding resource leaks. - Directly wrote the content returned by
getText()instead of usingprint(), which eliminates any unexpected extra formatting thatprint()might add.
A quick note: The python-docx library handles the DOCX file's internal encoding automatically, so you don't need to worry about decoding the file yourself—just pass the filename to docx.Document() and it takes care of the rest.
If you still run into edge cases with super rare special characters, you can add an error-handling fallback when writing:
with open(out_filename, 'w', encoding='utf8', errors='replace') as outfile: outfile.write(doc_content)
The errors='replace' flag will swap any unencodable characters with � instead of throwing an error, which can be a helpful safety net for non-standard DOCX files.
内容的提问来源于stack exchange,提问作者Rowland

