使用PyPDF2提取PDF文本时字符间多余空格的解决办法
问题:PyPDF2提取PDF文本时单词内出现多余空格的解决方法
我正在处理PDF文件,使用PyPDF2进行文本提取,提取过程中遇到同一单词字符间出现多余空格的问题,例如:
le vels应为levelsgood-qua lity应为good-quality
代码示例
from PyPDF2 import PdfReader reader = PdfReader("00001926B.pdf") page = reader.pages[80] text = page.extract_text() print(text)
提取的输出结果
2015 Microchip Technology Inc. DS00001926B-page 81LAN9354 9.2.2.6 100M Phase Lock Loop (PLL) The 100M PLL locks onto the reference clock and generates the 125 MHz clock used to drive the 125 MHz logic and the 100BASE-TX Transmitter. 9.2.3 100BASE-TX RECEIVE The 100BASE-TX receive data path is shown in Figure 9-3 . Shaded blocks are those which are internal to the PHY. Each major block is explained in the following sections. 9.2.3.1 100M Receive Input The MLT-3 data from the cable is fed into the PHY on inputs RXPx and RXNx via a 1:1 ratio transformer. The ADC sam- ples the incoming differential signal at a rate of 125M sa mples per second. Using a 64-level quantizer, 6 digital bits are generated to represent each sample. The DSP adjusts the gain of the ADC according to the observed signal levels suchthat the full dynamic range of the ADC can be used. 9.2.3.2 Equalizer, BLW Correction and Clock/Data Recovery The 6 bits from the ADC are fed into the DSP block. The equalizer in the DSP section compensates for phase and ampli- tude distortion caused by the physical channel consisting of magnetics, connectors, and CAT- 5 cable. The equalizer can restore the signal for any good-qua lity CAT-5 cable between 1m and 100m. If the DC content of t he signal is such that the low-frequency comp onents fall below the low frequency pole of the iso- lation transformer, then the droop characteristics of the transformer will become significant and Baseline Wander (BLW) on the received signal will result. To prevent corruption of the received data, the PHY corrects for BLW and can receive the ANSI X3.263-1995 FDDI TP-PMD defined “killer packet” with no bit errors. The 100M PLL generates multiple phases of the 125MHz clock. A multip lexer, controlled by the timing unit of the DSP, selects the optimum phase for sampling the data. This is used as the received recovered clock. This clock is used to extract the serial data from the received signal. 9.2.3.3 NRZI and MLT-3 Decoding The DSP generates the MLT-3 recovered le vels that are fed to the MLT-3 converter. The MLT-3 is then converted to an NRZI data stream.FIGURE 9-3: 100BASE-TX RECEIVE DATA PATH Port x MAC A/D ConverterMLT-3 ConverterNRZI Converter4B/5B Decoder Magnetics CAT-5 RJ45100M PLL Internal MII 25MHz by 4 bitsInternal MII Receive Clock 25MHz by 5 bits NRZI MLT-3 MLT-3 MLT-3 6 bit DataDescrambler and SIPO 125 Mbps Serial DSP: Timing recovery, Equalizer and BLW CorrectionMLT-3MII MAC Interface25MHz by 4 bits
解决方法
1. 正则表达式后处理
这类空格大多是PDF中单词跨换行拆分导致的,可以用正则匹配并合并:
import re # 处理普通单词拆分(如 le vels → levels) fixed_text = re.sub(r'(\w)\s+(\w)', r'\1\2', text) # 处理带连字符的拆分(如 good-qua lity → good-quality) fixed_text = re.sub(r'(\w-)\s+(\w)', r'\1\2', fixed_text) print(fixed_text)
注意:该方法可能会误删正常空格(比如两个短单词之间的空格),建议提取后人工抽查关键段落。
2. 换用更精准的PDF提取库
pdfplumber对文本布局的解析更智能,能更好识别跨换行的单词,减少这类错误:
import pdfplumber with pdfplumber.open("00001926B.pdf") as pdf: page = pdf.pages[80] text = page.extract_text() print(text)
3. 拼写检查辅助修正
使用pyenchant库进行拼写检查,自动修正拆分的单词(需先安装库和对应英文词典):
import enchant from enchant.checker import SpellChecker # 初始化英文拼写检查器 chkr = SpellChecker("en_US") chkr.set_text(text) # 自动修正拆分的单词 for err in chkr: # 仅处理单词内有空格的情况 if ' ' in err.word: corrected = err.word.replace(' ', '') if chkr.dict.check(corrected): err.replace(corrected) fixed_text = chkr.get_text() print(fixed_text)
内容的提问来源于stack exchange,提问作者Muhammad Samadzade
相关产品推荐
相关产品推荐

