Ghostscript生成PDF/UA合规文件在Firefox中无障碍失效,如何让标签字典内联?
Ghostscript生成PDF/UA无障碍标签兼容Firefox的解决方案
问题背景
通过Ghostscript从PostScript文件生成符合PDF/UA和WCAG 2.0标准的PDF时,用/BDC pdfmark实现的无障碍特性在Acrobat Reader搭配NVDA屏幕阅读器时可正常读取,但在Firefox中打开时,标题、表格、列表等标签的无障碍特性完全丢失。经排查,问题出在Ghostscript对标签关联字典的处理方式:它将标签字典转为页面资源引用(比如/Rnn指向Resources->Properties字典),而非直接内联在标签命令中。
可行解决方法
目前Ghostscript没有直接的配置项支持将标签字典内联,但可以通过以下两种方式实现需求:
方法1:自定义PostScript补丁逻辑
在PostScript代码中重定义/BDC的处理逻辑,拦截默认的资源引用生成行为,直接将字典内容嵌入到内容流里。示例代码如下:
% 重定义BDC命令处理 /BDC { dup type /arraytype eq { dup length 3 eq { 0 get dup /H1 eq { % 针对H1标签,可扩展到H2、Table、LI等其他标签 exch 1 get % 获取<< /MCID 0 >>字典 % 直接输出内联格式的BDC命令 cvs ( ) print ( << ) print { 0 get cvs print ( ) print 1 get cvs print ( ) print } forall ( >> BDC ) print pop pop pop exit } if } if } if % 执行原BDC命令逻辑 systemdict /BDC get exec } bind def % 原有标签代码保持不变 [ /H1 << /MCID 0 >> /BDC pdfmark (Heading 1) 7747 2920 mvs [ /EMC pdfmark
方法2:PDF生成后批量修改内容流
用PDF处理工具(比如PyPDF2、pdftk)提取页面资源中的Properties字典内容,批量替换内容流里的/Rnn引用为实际字典。示例Python代码(依赖PyPDF2库):
from PyPDF2 import PdfReader, PdfWriter reader = PdfReader("input.pdf") writer = PdfWriter() for page in reader.pages: # 获取页面资源的Properties字典 props = page.get("/Resources", {}).get("/Properties", {}) if not props: writer.add_page(page) continue # 读取并解码内容流 content = page.get_contents().decode() # 替换所有/Rnn引用为实际字典 for ref_name, obj_ref in props.items(): if ref_name.startswith("/R"): obj = reader.get_object(obj_ref) # 把PDF对象转为字符串格式的字典 dict_parts = [] for k, v in obj.items(): dict_parts.append(f"/{k} {v}") dict_str = f"<< {' '.join(dict_parts)} >>" # 替换内容流中的引用 content = content.replace(f"{ref_name} BDC", f"{ref_name} {dict_str} BDC") # 更新页面内容流 page.set_contents(content.encode()) writer.add_page(page) with open("output.pdf", "wb") as f: writer.write(f)
注意事项
- 方法1需要熟悉PostScript语法,需针对每种标签类型逐一适配逻辑。
- 方法2需确保原PDF结构规范,避免修改时破坏其他无障碍相关属性。
- 修改后需同时在Acrobat Reader和Firefox中测试,确保两边都能正常识别标签。
内容的提问来源于stack exchange,提问作者Giovanni
相关产品推荐
相关产品推荐

