You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Ghostscript生成PDF/UA合规文件在Firefox中无障碍失效,如何让标签字典内联?

Ghostscript生成PDF/UA无障碍标签兼容Firefox的解决方案

问题背景

通过Ghostscript从PostScript文件生成符合PDF/UA和WCAG 2.0标准的PDF时,用/BDC pdfmark实现的无障碍特性在Acrobat Reader搭配NVDA屏幕阅读器时可正常读取,但在Firefox中打开时,标题、表格、列表等标签的无障碍特性完全丢失。经排查,问题出在Ghostscript对标签关联字典的处理方式:它将标签字典转为页面资源引用(比如/Rnn指向Resources->Properties字典),而非直接内联在标签命令中。

可行解决方法

目前Ghostscript没有直接的配置项支持将标签字典内联,但可以通过以下两种方式实现需求:

方法1:自定义PostScript补丁逻辑

在PostScript代码中重定义/BDC的处理逻辑,拦截默认的资源引用生成行为,直接将字典内容嵌入到内容流里。示例代码如下:

% 重定义BDC命令处理
/BDC {
  dup type /arraytype eq {
    dup length 3 eq {
      0 get dup /H1 eq { % 针对H1标签,可扩展到H2、Table、LI等其他标签
        exch 1 get % 获取<< /MCID 0 >>字典
        % 直接输出内联格式的BDC命令
        cvs ( ) print
        ( << ) print
        {
          0 get cvs print
          ( ) print
          1 get cvs print
          ( ) print
        } forall
        ( >> BDC ) print
        pop pop pop
        exit
      } if
    } if
  } if
  % 执行原BDC命令逻辑
  systemdict /BDC get exec
} bind def

% 原有标签代码保持不变
[ /H1 << /MCID 0 >> /BDC pdfmark
(Heading 1) 7747 2920 mvs
[ /EMC pdfmark

方法2:PDF生成后批量修改内容流

用PDF处理工具(比如PyPDF2、pdftk)提取页面资源中的Properties字典内容,批量替换内容流里的/Rnn引用为实际字典。示例Python代码(依赖PyPDF2库):

from PyPDF2 import PdfReader, PdfWriter

reader = PdfReader("input.pdf")
writer = PdfWriter()

for page in reader.pages:
    # 获取页面资源的Properties字典
    props = page.get("/Resources", {}).get("/Properties", {})
    if not props:
        writer.add_page(page)
        continue
    # 读取并解码内容流
    content = page.get_contents().decode()
    # 替换所有/Rnn引用为实际字典
    for ref_name, obj_ref in props.items():
        if ref_name.startswith("/R"):
            obj = reader.get_object(obj_ref)
            # 把PDF对象转为字符串格式的字典
            dict_parts = []
            for k, v in obj.items():
                dict_parts.append(f"/{k} {v}")
            dict_str = f"<< {' '.join(dict_parts)} >>"
            # 替换内容流中的引用
            content = content.replace(f"{ref_name} BDC", f"{ref_name} {dict_str} BDC")
    # 更新页面内容流
    page.set_contents(content.encode())
    writer.add_page(page)

with open("output.pdf", "wb") as f:
    writer.write(f)

注意事项

  • 方法1需要熟悉PostScript语法,需针对每种标签类型逐一适配逻辑。
  • 方法2需确保原PDF结构规范,避免修改时破坏其他无障碍相关属性。
  • 修改后需同时在Acrobat Reader和Firefox中测试,确保两边都能正常识别标签。

内容的提问来源于stack exchange,提问作者Giovanni

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 11:22:48