使用PyPDF2与NLTK构建PDF语料库时遇UnicodeDecodeError求助
解决UnicodeDecodeError:PyPDF2提取文本与NLTK语料库读取的编码问题
咱们先拆解一下这个错误的根源:UnicodeDecodeError: 'utf-8' codec can't decode byte 0x80 in position 395: invalid start byte 说明NLTK在读取你生成的txt文件时,遇到了无法用UTF-8解码的字符。这大概率是两个环节出了问题:
- PyPDF2从PDF提取的文本本身包含非UTF-8编码的特殊字符
- 你保存txt文件时没有正确处理编码,或者NLTK读取时没有指定匹配的编码
下面是具体的解决步骤:
一、修正文本保存时的编码处理
当你把PDF提取的文本写入txt文件时,需要显式指定UTF-8编码,同时处理那些无法用UTF-8编码的特殊字符(比如用errors='replace'替换成占位符,或者errors='ignore'忽略)。
修改你保存文件的代码块:
for idx, f in enumerate(files): # 显式指定编码为utf-8,并处理异常字符 with open(corpusDir + str(idx) + '.txt', 'w', encoding='utf-8', errors='replace') as output: output.write(f)
二、让NLTK读取语料库时指定编码
NLTK的PlaintextCorpusReader默认会用系统默认编码读取文件,咱们需要显式指定UTF-8编码,确保和保存时的编码一致:
修改初始化corpus的代码:
# 指定encoding='utf-8',和保存文件时的编码匹配 corpus = PlaintextCorpusReader(corpusDir, '.*', encoding='utf-8')
三、可选:优化PDF文本提取的编码处理
如果上面的步骤还没解决问题,可能是PyPDF2提取的文本本身编码混乱,咱们可以在提取文本后先做一次编码转换的处理:
修改getTextPDF函数:
def getTextPDF(pdfFileName): pdf_file = open(pdfFileName, 'rb') readpdf = PdfFileReader(pdf_file) text = [] for i in range(readpdf.getNumPages()): page_text = readpdf.getPage(i).extractText() # 先将文本转为字节再解码,处理潜在的编码问题 try: page_text = page_text.encode('utf-8', errors='replace').decode('utf-8') except: pass text.append(page_text) return '\n'.join(text)
修改后的完整代码
import nltk import PyPDF2 from nltk.corpus.reader.plaintext import PlaintextCorpusReader from PyPDF2 import PdfFileReader def getTextPDF(pdfFileName): pdf_file = open(pdfFileName, 'rb') readpdf = PdfFileReader(pdf_file) text = [] for i in range(readpdf.getNumPages()): page_text = readpdf.getPage(i).extractText() # 处理提取文本的编码异常 try: page_text = page_text.encode('utf-8', errors='replace').decode('utf-8') except: pass text.append(page_text) return '\n'.join(text) corpusDir = 'reports/' jun15 = getTextPDF('reports/June2015.pdf') dec15 = getTextPDF('reports/December2015.pdf') jun16 = getTextPDF('reports/June2016.pdf') dec16 = getTextPDF('reports/December2016.pdf') jun17 = getTextPDF('reports/June2017.pdf') dec17 = getTextPDF('reports/December2017.pdf') files = [jun15,dec15,jun16,dec16,jun17,dec17] for idx, f in enumerate(files): # 保存时指定UTF-8编码并处理异常字符 with open(corpusDir + str(idx) + '.txt', 'w', encoding='utf-8', errors='replace') as output: output.write(f) # 读取语料库时指定UTF-8编码 corpus = PlaintextCorpusReader(corpusDir, '.*', encoding='utf-8') print(corpus.words())
关键说明
errors='replace'的作用:当遇到无法编码/解码的字符时,会用�替换,避免程序崩溃,同时保留大部分文本内容。如果不需要保留这些异常字符,可以换成errors='ignore'直接忽略。- 显式指定编码的重要性:不管是保存文件还是读取文件,统一用UTF-8编码可以避免绝大多数编码不匹配的问题。
内容的提问来源于stack exchange,提问作者user9221735
相关产品推荐
相关产品推荐

