You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实现docx转纯文本代码触发UnicodeDecodeError问题求助

Fixing UnicodeDecodeError When Converting DOCX to TXT in Python

Hey there, let's break down what's going wrong with your code and how to fix that frustrating UnicodeDecodeError.

First, let's pinpoint the root issue in your original code:

You’re using open(file) to read the DOCX file directly in your loop, but this is completely unnecessary—docx.Document(filename) already handles loading the DOCX file properly behind the scenes. When you use open(file) without specifying an encoding, Python defaults to your system’s native charmap encoding (often limited on Windows), which can’t parse certain special characters in your DOCX file, triggering the decode error.

Here's the cleaned-up, error-free version of your code, with key improvements explained:

import docx
import glob

def getText(filename):
    doc = docx.Document(filename)
    fullText = []
    for para in doc.paragraphs:
        fullText.append(para.text)
    return '\n'.join(fullText)

# Loop through all DOCX files in the current folder
for file in glob.glob('*.docx'):
    out_filename = file.replace('.docx', '-out.txt')
    # Fetch the processed text content directly from your function
    doc_content = getText(file)
    # Use a with-statement to handle file writing (auto-closes the file)
    with open(out_filename, 'w', encoding='utf8') as outfile:
        outfile.write(doc_content)

Let's walk through the critical fixes:

  • Removed the redundant for line in open(file): loop—this was the main cause of your encoding error, since we never need to read a DOCX file as raw plain text.
  • Switched to a with statement for writing the output file: this is Python best practice, as it automatically closes the file after writing, avoiding resource leaks.
  • Directly wrote the content returned by getText() instead of using print(), which eliminates any unexpected extra formatting that print() might add.

A quick note: The python-docx library handles the DOCX file's internal encoding automatically, so you don't need to worry about decoding the file yourself—just pass the filename to docx.Document() and it takes care of the rest.

If you still run into edge cases with super rare special characters, you can add an error-handling fallback when writing:

with open(out_filename, 'w', encoding='utf8', errors='replace') as outfile:
    outfile.write(doc_content)

The errors='replace' flag will swap any unencodable characters with � instead of throwing an error, which can be a helpful safety net for non-standard DOCX files.

内容的提问来源于stack exchange,提问作者Rowland

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:03:57