You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pdfplumber读取PDF时单页旋转引发全页旋转值异常问题

解决pdfplumber读取PDF页面旋转值错误的问题

你的问题是pdfplumber返回的页面旋转值不符合预期:第一页实际无需旋转(应为0),但输出和第二页一样是270。这通常是因为pdfplumber的page.rotation会叠加文档级全局旋转和页面自身旋转,或者部分PDF的文本旋转是通过内容变换矩阵(CTM)实现而非页面的/Rotate字段导致的。

解决方案1:读取页面原始旋转值(排除文档级干扰)

pdfplumber中page.rotation的计算逻辑是(文档旋转 + 页面自身旋转) % 360,如果文档本身设置了全局旋转,会影响所有页面的结果。可以直接读取页面字典中的原始/Rotate字段:

import pdfplumber

pdf_path = "some_pdf.pdf"

with pdfplumber.open(pdf_path) as pdf:
    for page_num, page in enumerate(pdf.pages, 1):
        # 获取页面自身的原始旋转值(不受文档全局旋转影响)
        raw_rotation = page.raw_page.get("/Rotate", 0)
        # 若需要实际显示的旋转值,可结合文档级旋转计算
        actual_rotation = (pdf.metadata.get("rotation", 0) + raw_rotation) % 360
        print(f"页面{page_num} - 原始旋转值: {raw_rotation}, 实际显示旋转值: {actual_rotation}")

解决方案2:从文本变换矩阵提取实际旋转(应对CTM旋转场景)

部分PDF不会设置页面的/Rotate字段,而是通过内容变换矩阵(CTM)旋转文本。此时需要从文本对象的变换信息中提取旋转角度:

import pdfplumber

pdf_path = "some_pdf.pdf"

def get_text_rotation(page):
    text_objects = page.extract_words()
    if not text_objects:
        # 无文本时 fallback 到页面原始旋转值
        return page.raw_page.get("/Rotate", 0)
    # 从第一个文本对象的变换矩阵判断旋转角度(仅适配90/180/270度标准旋转)
    ctm = text_objects[0]["transform"]
    if ctm[0] == 0 and ctm[1] == -1:
        return 90
    elif ctm[0] == -1 and ctm[1] == 0:
        return 180
    elif ctm[0] == 0 and ctm[1] == 1:
        return 270
    else:
        return 0

with pdfplumber.open(pdf_path) as pdf:
    for page_num, page in enumerate(pdf.pages, 1):
        rotation = get_text_rotation(page)
        print(f"页面{page_num}实际文本旋转值: {rotation}")

验证与说明

  • 先运行方案1,若输出的原始旋转值第一页为0、第二页为270,说明问题出在文档级全局旋转叠加,此时用raw_rotation即可得到正确的页面自身旋转值。
  • 若方案1输出仍不符合预期,说明PDF用了CTM旋转文本,改用方案2即可获取文本实际的旋转角度。

内容的提问来源于stack exchange,提问作者jsiller

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 13:45:24