使用pdfplumber读取PDF时单页旋转引发全页旋转值异常问题
解决pdfplumber读取PDF页面旋转值错误的问题
你的问题是pdfplumber返回的页面旋转值不符合预期:第一页实际无需旋转(应为0),但输出和第二页一样是270。这通常是因为pdfplumber的page.rotation会叠加文档级全局旋转和页面自身旋转,或者部分PDF的文本旋转是通过内容变换矩阵(CTM)实现而非页面的/Rotate字段导致的。
解决方案1:读取页面原始旋转值(排除文档级干扰)
pdfplumber中page.rotation的计算逻辑是(文档旋转 + 页面自身旋转) % 360,如果文档本身设置了全局旋转,会影响所有页面的结果。可以直接读取页面字典中的原始/Rotate字段:
import pdfplumber pdf_path = "some_pdf.pdf" with pdfplumber.open(pdf_path) as pdf: for page_num, page in enumerate(pdf.pages, 1): # 获取页面自身的原始旋转值(不受文档全局旋转影响) raw_rotation = page.raw_page.get("/Rotate", 0) # 若需要实际显示的旋转值,可结合文档级旋转计算 actual_rotation = (pdf.metadata.get("rotation", 0) + raw_rotation) % 360 print(f"页面{page_num} - 原始旋转值: {raw_rotation}, 实际显示旋转值: {actual_rotation}")
解决方案2:从文本变换矩阵提取实际旋转(应对CTM旋转场景)
部分PDF不会设置页面的/Rotate字段,而是通过内容变换矩阵(CTM)旋转文本。此时需要从文本对象的变换信息中提取旋转角度:
import pdfplumber pdf_path = "some_pdf.pdf" def get_text_rotation(page): text_objects = page.extract_words() if not text_objects: # 无文本时 fallback 到页面原始旋转值 return page.raw_page.get("/Rotate", 0) # 从第一个文本对象的变换矩阵判断旋转角度(仅适配90/180/270度标准旋转) ctm = text_objects[0]["transform"] if ctm[0] == 0 and ctm[1] == -1: return 90 elif ctm[0] == -1 and ctm[1] == 0: return 180 elif ctm[0] == 0 and ctm[1] == 1: return 270 else: return 0 with pdfplumber.open(pdf_path) as pdf: for page_num, page in enumerate(pdf.pages, 1): rotation = get_text_rotation(page) print(f"页面{page_num}实际文本旋转值: {rotation}")
验证与说明
- 先运行方案1,若输出的
原始旋转值第一页为0、第二页为270,说明问题出在文档级全局旋转叠加,此时用raw_rotation即可得到正确的页面自身旋转值。 - 若方案1输出仍不符合预期,说明PDF用了CTM旋转文本,改用方案2即可获取文本实际的旋转角度。
内容的提问来源于stack exchange,提问作者jsiller
相关产品推荐
相关产品推荐

