You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python提取PDF中粗体文本后的语句?

提取PDF粗体文本后的内容解决方案

原代码只筛选了粗体字符,所以只能拿到粗体部分。要获取粗体之后的内容,得先定位粗体在字符序列里的位置,再提取后续内容。

实现思路

  • 先获取页面的所有字符对象,保留每个字符的字体信息和顺序
  • 找到最后一个粗体字符的位置
  • 从该位置之后的字符开始,提取对应的文本

修改后的代码

import pdfplumber

with pdfplumber.open('C:/Users/somefile.pdf') as pdf:
    for page_num in range(12, 15):
        page = pdf.pages[page_num]
        # 获取页面所有字符对象,按默认顺序排列
        all_chars = page.chars
        # 收集所有粗体字符的索引
        bold_indices = []
        for idx, char in enumerate(all_chars):
            if "Bold" in char["fontname"]:
                bold_indices.append(idx)
        
        if not bold_indices:
            # 当前页面无粗体内容,直接跳过
            continue
        
        # 定位最后一个粗体字符的位置
        last_bold_idx = bold_indices[-1]
        # 提取最后一个粗体之后的所有字符
        target_chars = all_chars[last_bold_idx + 1:]
        # 拼接成完整文本并去除首尾空格
        target_text = ''.join([char["text"] for char in target_chars]).strip()
        print(f"第{page_num+1}页粗体后内容:\n{target_text}\n")

补充优化

  • 如果只想提取粗体之后的第一句话,可以在得到target_text后做分割处理:
    # 以句号为分隔符取第一句,可根据需求替换成感叹号、问号等
    first_sentence = target_text.split('.')[0] + '.' if '.' in target_text else target_text
    print(f"第{page_num+1}页粗体后第一句:{first_sentence}")
    
  • 若遇到PDF排版导致字符顺序错乱的情况,可以结合字符的坐标重新排序:
    # 按纵向(top)和横向(x0)坐标排序,确保字符顺序符合阅读逻辑
    sorted_chars = sorted(all_chars, key=lambda c: (c["top"], c["x0"]))
    

内容的提问来源于stack exchange,提问作者Ben

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 18:28:01