You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用PyPDF2与NLTK构建PDF语料库时遇UnicodeDecodeError求助

解决UnicodeDecodeError:PyPDF2提取文本与NLTK语料库读取的编码问题

咱们先拆解一下这个错误的根源:UnicodeDecodeError: 'utf-8' codec can't decode byte 0x80 in position 395: invalid start byte 说明NLTK在读取你生成的txt文件时,遇到了无法用UTF-8解码的字符。这大概率是两个环节出了问题:

  • PyPDF2从PDF提取的文本本身包含非UTF-8编码的特殊字符
  • 你保存txt文件时没有正确处理编码,或者NLTK读取时没有指定匹配的编码

下面是具体的解决步骤:

一、修正文本保存时的编码处理

当你把PDF提取的文本写入txt文件时,需要显式指定UTF-8编码,同时处理那些无法用UTF-8编码的特殊字符(比如用errors='replace'替换成占位符,或者errors='ignore'忽略)。

修改你保存文件的代码块:

for idx, f in enumerate(files):
    # 显式指定编码为utf-8,并处理异常字符
    with open(corpusDir + str(idx) + '.txt', 'w', encoding='utf-8', errors='replace') as output:
        output.write(f)

二、让NLTK读取语料库时指定编码

NLTK的PlaintextCorpusReader默认会用系统默认编码读取文件,咱们需要显式指定UTF-8编码,确保和保存时的编码一致:

修改初始化corpus的代码:

# 指定encoding='utf-8',和保存文件时的编码匹配
corpus = PlaintextCorpusReader(corpusDir, '.*', encoding='utf-8')

三、可选:优化PDF文本提取的编码处理

如果上面的步骤还没解决问题,可能是PyPDF2提取的文本本身编码混乱,咱们可以在提取文本后先做一次编码转换的处理:

修改getTextPDF函数:

def getTextPDF(pdfFileName):
    pdf_file = open(pdfFileName, 'rb')
    readpdf = PdfFileReader(pdf_file)
    text = []
    for i in range(readpdf.getNumPages()):
        page_text = readpdf.getPage(i).extractText()
        # 先将文本转为字节再解码,处理潜在的编码问题
        try:
            page_text = page_text.encode('utf-8', errors='replace').decode('utf-8')
        except:
            pass
        text.append(page_text)
    return '\n'.join(text)

修改后的完整代码

import nltk
import PyPDF2
from nltk.corpus.reader.plaintext import PlaintextCorpusReader
from PyPDF2 import PdfFileReader

def getTextPDF(pdfFileName):
    pdf_file = open(pdfFileName, 'rb')
    readpdf = PdfFileReader(pdf_file)
    text = []
    for i in range(readpdf.getNumPages()):
        page_text = readpdf.getPage(i).extractText()
        # 处理提取文本的编码异常
        try:
            page_text = page_text.encode('utf-8', errors='replace').decode('utf-8')
        except:
            pass
        text.append(page_text)
    return '\n'.join(text)

corpusDir = 'reports/'
jun15 = getTextPDF('reports/June2015.pdf')
dec15 = getTextPDF('reports/December2015.pdf')
jun16 = getTextPDF('reports/June2016.pdf')
dec16 = getTextPDF('reports/December2016.pdf')
jun17 = getTextPDF('reports/June2017.pdf')
dec17 = getTextPDF('reports/December2017.pdf')

files = [jun15,dec15,jun16,dec16,jun17,dec17]

for idx, f in enumerate(files):
    # 保存时指定UTF-8编码并处理异常字符
    with open(corpusDir + str(idx) + '.txt', 'w', encoding='utf-8', errors='replace') as output:
        output.write(f)

# 读取语料库时指定UTF-8编码
corpus = PlaintextCorpusReader(corpusDir, '.*', encoding='utf-8')
print(corpus.words())

关键说明

  • errors='replace'的作用:当遇到无法编码/解码的字符时,会用�替换,避免程序崩溃,同时保留大部分文本内容。如果不需要保留这些异常字符,可以换成errors='ignore'直接忽略。
  • 显式指定编码的重要性:不管是保存文件还是读取文件,统一用UTF-8编码可以避免绝大多数编码不匹配的问题。

内容的提问来源于stack exchange,提问作者user9221735

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 08:48:41