You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

不同严格度PDF对比:忽略元数据的视觉一致性验证需求

问题:验证PDF视觉一致性(忽略元数据差异)

我有两个文件夹,各包含约100个由同一PDF生成程序不同运行实例生成的PDF文件。修改程序后,生成的PDF应在布局、字体、图表等视觉层面完全一致,仅需忽略因运行时间不同导致的元数据差异。


尝试过的方法

1. 文件哈希值对比

最初尝试直接对比文件哈希值,代码如下:

h1 = hashlib.sha1()
h2 = hashlib.sha1()

with open(fileName1, "rb") as file:
    chunk = 0
    while chunk != b'':
        chunk = file.read(1024)
        h1.update(chunk)

with open(fileName2, "rb") as file:
    chunk = 0
    while chunk != b'':
        chunk = file.read(1024)
        h2.update(chunk)

return (h1.hexdigest() == h2.hexdigest())

但该方法始终返回False,推测是时间相关元数据导致差异。

2. 重置元数据后再哈希

尝试将PDF的修改时间和创建时间设为None,再做哈希对比:

pdf1 = pdfrw.PdfReader(fileName1)
pdf1.Info.ModDate = pdf1.Info.CreationDate = None
pdfrw.PdfWriter().write(fileName1, pdf1)
    
pdf2 = pdfrw.PdfReader(fileName2)
pdf2.Info.ModDate = pdf2.Info.CreationDate = None
pdfrw.PdfWriter().write(fileName2, pdf2)

遍历文件夹文件执行该操作后再哈希,结果有时为True有时为False,不稳定。

3. XRef(交叉引用表)对比

在他人帮助下,尝试了两种基于XRef的对比方法:

完整XRef对比

doc1 = fitz.open(fileName1)
xrefs1 = doc1.xref_length() # cross reference table 1
doc2 = fitz.open(fileName2)
xrefs2 = doc2.xref_length() # cross reference table 2
    
if (xrefs1 != xrefs2):
    print("Files are not equal")
    return False
    
for xref in range(1, xrefs1):  # loop over objects, index 0 must be skipped
    # compare the PDF object definition sources
    if (doc1.xref_object(xref) != doc2.xref_object(xref)):
        print(f"Files differ at xref {xref}.")
        return False
    if doc1.xref_is_stream(xref):  # compare binary streams
        stream1 = doc1.xref_stream_raw(xref)  # read binary stream
        try:
            stream2 = doc2.xref_stream_raw(xref)  # read binary stream
        except:  # stream extraction doc2 did not work!
            print(f"stream discrepancy at xref {xref}")
            return False
        if (stream1 != stream2):
            print(f"stream discrepancy at xref {xref}")
            return False
return True

忽略元数据的XRef对比

doc1 = fitz.open(fileName1)
xrefs1 = doc1.xref_length() # cross reference table 1
doc2 = fitz.open(fileName2)
xrefs2 = doc2.xref_length() # cross reference table 2
    
info1 = doc1.xref_get_key(-1, "Info")  # extract the info object
info2 = doc2.xref_get_key(-1, "Info")
    
if (info1 != info2):
    print("Unequal info objects")
    return False
    
if (info1[0] == "xref"): # is there metadata at all?
    info_xref1 = int(info1[1].split()[0])  # xref of info object doc1
    info_xref2 = int(info2[1].split()[0])  # xref of info object doc1

else:
    info_xref1 = 0
            
for xref in range(1, xrefs1):  # loop over objects, index 0 must be skipped
    # compare the PDF object definition sources
    if (xref != info_xref1):
        if (doc1.xref_object(xref) != doc2.xref_object(xref)):
            print(f"Files differ at xref {xref}.")
            return False
        if doc1.xref_is_stream(xref):  # compare binary streams
            stream1 = doc1.xref_stream_raw(xref)  # read binary stream
            try:
                stream2 = doc2.xref_stream_raw(xref)  # read binary stream
            except:  # stream extraction doc2 did not work!
                print(f"stream discrepancy at xref {xref}")
                return False
            if (stream1 != stream2):
                print(f"stream discrepancy at xref {xref}")
                return False
return True

但对已重置时间戳的PDF执行这两个方法时,部分文件返回True,部分返回False,依然无法稳定验证视觉一致性。


需求

我使用reportlab库生成PDF,想知道是否存在无需将页面导出为图片的方法,就能验证PDF的视觉一致性——即忽略内部结构差异,只确认布局、字体、图表等视觉表现完全一致。

内容的提问来源于stack exchange,提问作者Hagbard

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 00:00:57