不同严格度PDF对比:忽略元数据的视觉一致性验证需求
问题:验证PDF视觉一致性(忽略元数据差异)
我有两个文件夹,各包含约100个由同一PDF生成程序不同运行实例生成的PDF文件。修改程序后,生成的PDF应在布局、字体、图表等视觉层面完全一致,仅需忽略因运行时间不同导致的元数据差异。
尝试过的方法
1. 文件哈希值对比
最初尝试直接对比文件哈希值,代码如下:
h1 = hashlib.sha1() h2 = hashlib.sha1() with open(fileName1, "rb") as file: chunk = 0 while chunk != b'': chunk = file.read(1024) h1.update(chunk) with open(fileName2, "rb") as file: chunk = 0 while chunk != b'': chunk = file.read(1024) h2.update(chunk) return (h1.hexdigest() == h2.hexdigest())
但该方法始终返回False,推测是时间相关元数据导致差异。
2. 重置元数据后再哈希
尝试将PDF的修改时间和创建时间设为None,再做哈希对比:
pdf1 = pdfrw.PdfReader(fileName1) pdf1.Info.ModDate = pdf1.Info.CreationDate = None pdfrw.PdfWriter().write(fileName1, pdf1) pdf2 = pdfrw.PdfReader(fileName2) pdf2.Info.ModDate = pdf2.Info.CreationDate = None pdfrw.PdfWriter().write(fileName2, pdf2)
遍历文件夹文件执行该操作后再哈希,结果有时为True有时为False,不稳定。
3. XRef(交叉引用表)对比
在他人帮助下,尝试了两种基于XRef的对比方法:
完整XRef对比
doc1 = fitz.open(fileName1) xrefs1 = doc1.xref_length() # cross reference table 1 doc2 = fitz.open(fileName2) xrefs2 = doc2.xref_length() # cross reference table 2 if (xrefs1 != xrefs2): print("Files are not equal") return False for xref in range(1, xrefs1): # loop over objects, index 0 must be skipped # compare the PDF object definition sources if (doc1.xref_object(xref) != doc2.xref_object(xref)): print(f"Files differ at xref {xref}.") return False if doc1.xref_is_stream(xref): # compare binary streams stream1 = doc1.xref_stream_raw(xref) # read binary stream try: stream2 = doc2.xref_stream_raw(xref) # read binary stream except: # stream extraction doc2 did not work! print(f"stream discrepancy at xref {xref}") return False if (stream1 != stream2): print(f"stream discrepancy at xref {xref}") return False return True
忽略元数据的XRef对比
doc1 = fitz.open(fileName1) xrefs1 = doc1.xref_length() # cross reference table 1 doc2 = fitz.open(fileName2) xrefs2 = doc2.xref_length() # cross reference table 2 info1 = doc1.xref_get_key(-1, "Info") # extract the info object info2 = doc2.xref_get_key(-1, "Info") if (info1 != info2): print("Unequal info objects") return False if (info1[0] == "xref"): # is there metadata at all? info_xref1 = int(info1[1].split()[0]) # xref of info object doc1 info_xref2 = int(info2[1].split()[0]) # xref of info object doc1 else: info_xref1 = 0 for xref in range(1, xrefs1): # loop over objects, index 0 must be skipped # compare the PDF object definition sources if (xref != info_xref1): if (doc1.xref_object(xref) != doc2.xref_object(xref)): print(f"Files differ at xref {xref}.") return False if doc1.xref_is_stream(xref): # compare binary streams stream1 = doc1.xref_stream_raw(xref) # read binary stream try: stream2 = doc2.xref_stream_raw(xref) # read binary stream except: # stream extraction doc2 did not work! print(f"stream discrepancy at xref {xref}") return False if (stream1 != stream2): print(f"stream discrepancy at xref {xref}") return False return True
但对已重置时间戳的PDF执行这两个方法时,部分文件返回True,部分返回False,依然无法稳定验证视觉一致性。
需求
我使用reportlab库生成PDF,想知道是否存在无需将页面导出为图片的方法,就能验证PDF的视觉一致性——即忽略内部结构差异,只确认布局、字体、图表等视觉表现完全一致。
内容的提问来源于stack exchange,提问作者Hagbard
相关产品推荐
相关产品推荐

