You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python读取并更新PDF中的图片表单字段?

如何处理PDF中的图片按钮表单字段(读取与更新)

PDF模板中的图片上传字段通常被定义为按钮字段(/FT为/Btn),这类字段的图片内容存储在外观字典(/AP)中,而非普通字段的/V值,所以直接用update_page_form_field_values方法无法更新图片。以下是具体的读取和更新方案:

一、读取图片字段中的图片

需要解析按钮字段的外观字典,提取正常状态(/N)下的图片对象:

from pypdf import PdfReader

reader = PdfReader("form.pdf")
fields = reader.get_fields()

# 定位目标图片字段
photo_field = fields.get("photograph")
if photo_field:
    # 获取按钮外观字典
    ap_dict = photo_field.get("/AP")
    if ap_dict and "/N" in ap_dict:
        appearance_obj = ap_dict["/N"]
        # 检查是否为直接的图片XObject
        if "/Subtype" in appearance_obj and appearance_obj["/Subtype"] == "/Image":
            # 提取图片字节数据并保存
            image_bytes = appearance_obj.get_data()
            with open("extracted_photo.png", "wb") as f:
                f.write(image_bytes)
            print("图片提取成功")
        else:
            print("外观对象非直接图片,需进一步解析")
    else:
        print("该按钮字段未设置图片内容")

二、更新图片字段的内容

需要将新图片转换为PDF兼容的XObject,替换按钮字段外观字典中的对应内容:

from pypdf import PdfReader, PdfWriter
from pypdf.generic import DictionaryObject, NameObject, StreamObject
from PIL import Image
import io

def convert_image_to_pdf_xobject(image_path, pdf_writer):
    # 用PIL读取图片并转为字节流
    img = Image.open(image_path)
    img_buffer = io.BytesIO()
    img.save(img_buffer, format="PNG")
    img_buffer.seek(0)
    
    # 创建PDF图片流对象
    img_stream = StreamObject()
    img_stream.set_data(img_buffer.getvalue())
    # 设置图片必要属性
    img_stream.update({
        NameObject("/Type"): NameObject("/XObject"),
        NameObject("/Subtype"): NameObject("/Image"),
        NameObject("/Width"): img.width,
        NameObject("/Height"): img.height,
        NameObject("/ColorSpace"): NameObject("/DeviceRGB"),
        NameObject("/BitsPerComponent"): 8,
    })
    # 将图片对象添加到PDF writer并返回引用
    return pdf_writer._objects.add(img_stream)

# 读取原始PDF并初始化writer
reader = PdfReader("form.pdf")
writer = PdfWriter()
writer.append_pages_from_reader(reader)

# 定位目标字段
photo_field = reader.get_fields().get("photograph")
if photo_field:
    # 转换新图片为PDF XObject
    new_img_ref = convert_image_to_pdf_xobject("new_photo.png", writer)
    
    # 初始化外观字典(如果不存在)
    if "/AP" not in photo_field:
        photo_field["/AP"] = DictionaryObject()
    
    # 更新正常状态的外观为新图片
    photo_field["/AP"][NameObject("/N")] = new_img_ref
    
    # 更新writer中的字段
    writer.update_field("photograph", photo_field)

# 保存修改后的PDF
with open("updated_form.pdf", "wb") as f:
    writer.write(f)

注意事项

  • 需提前安装依赖:pip install pypdf pillow
  • 部分PDF的按钮字段可能包含多种状态(如按下/悬停),若需要同步更新,需修改外观字典中的/D(按下状态)、/R(悬停状态)等属性
  • 若按钮外观是复合XObject(包含多个元素),则需要更复杂的解析替换逻辑

内容的提问来源于stack exchange,提问作者Rahul

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 21:42:39