You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用PyMuPDF提取非结构化PDF数据时遇AttributeError问题求助

问题:PyMuPDF提取PDF数据时出现AttributeError错误

我正在按照《从非结构化PDF中提取数据》教程使用PyMuPDF提取非结构化PDF数据,作为Python新手,运行代码时出现如下错误:

AttributeError                            Traceback (most recent call last)
<ipython-input-2-7f394b979351> in <module>
      1 first_annots=[]
      2 
----> 3 rec=page1.first_annot.rect
      4 
      5 rec

AttributeError: 'NoneType' object has no attribute 'rect'

运行的代码如下:

import fitz
import pandas as pd 
doc = fitz.open('Mansfield--70-21009048 - ConvertToExcel.pdf')
page1 = doc[0]
words = page1.get_text("words")
words[0]

first_annots=[]

rec=page1.first_annot.rect

rec

#Information of words in first object is stored in mywords

mywords = [w for w in words if fitz.Rect(w[:4]) in rec]

ann= make_text(mywords)

first_annots.append(ann)

def make_text(words):

    line_dict = {} 

    words.sort(key=lambda w: w[0])

    for w in words:  

        y1 = round(w[3], 1)  

        word = w[4] 

        line = line_dict.get(y1, [])  

        line.append(word)  

        line_dict[y1] = line  

    lines = list(line_dict.items())

    lines.sort()  

    return "\n".join([" ".join(line[1]) for line in lines])

print(rec)
print(first_annots)

错误原因及解决方案

  1. 错误根源
    错误提示'NoneType' object has no attribute 'rect',说明page1.first_annot的值是None——你的目标PDF页面没有任何注释(Annotation,比如高亮、批注、框选标记等),而教程代码默认页面存在注释,直接调用first_annot.rect就会报错。

  2. 针对性解决

    • 情况1:PDF确实没有注释,想提取特定区域文本
      手动定义目标区域的坐标(坐标格式为(x0, y0, x1, y1),可通过PDF阅读器的坐标工具获取),替换原有的rec=page1.first_annot.rect:

      # 示例:定义一个左上角(100,100)到右下角(500,300)的矩形区域
      rec = fitz.Rect(100, 100, 500, 300)
      
    • 情况2:需要基于注释提取,但PDF无注释
      要么先给PDF添加注释,要么在代码里先判断注释是否存在,避免报错:

      first_annots = []
      # 先检查页面是否有注释
      if page1.first_annot:
          rec = page1.first_annot.rect
          mywords = [w for w in words if fitz.Rect(w[:4]) in rec]
          ann = make_text(mywords)
          first_annots.append(ann)
      else:
          print("当前页面未检测到注释")
      
    • 额外修复:函数定义顺序问题
      原代码中make_text函数定义在调用之后,会触发NameError,需要把函数定义移到调用之前:

      # 先定义make_text函数
      def make_text(words):
          line_dict = {} 
          words.sort(key=lambda w: w[0])
          for w in words:  
              y1 = round(w[3], 1)  
              word = w[4] 
              line = line_dict.get(y1, [])  
              line.append(word)  
              line_dict[y1] = line  
          lines = list(line_dict.items())
          lines.sort()  
          return "\n".join([" ".join(line[1]) for line in lines])
      
      # 再执行后续逻辑
      first_annots=[]
      # ... 其余代码 ...
      

内容的提问来源于stack exchange,提问作者Mech_Saran

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 00:45:58