HTML数学公式转Asciidoc:Pandoc转换失效的解决方法?
HTML数学公式转Asciidoc Stem块的解决方案
一、通过Pandoc预处理实现正确转换
Pandoc可以通过自定义Lua过滤器解决这个问题,核心是识别HTML中的数学公式元素,将其转换为Asciidoc标准的Stem块格式:
- 编写Lua过滤器:创建一个名为
math-convert.lua的文件,根据你的HTML公式标签结构编写匹配逻辑。比如针对MathML的<math>标签或带特定类的文本容器:-- 处理原生MathML元素 function Math(el) return pandoc.RawBlock("asciidoc", "[latexmath]\n++++\n" .. el.text .. "\n++++") end -- 处理带"math"类的span/div标签 function Span(el) if el.classes:includes("math") then return pandoc.RawBlock("asciidoc", "[latexmath]\n++++\n" .. el.text .. "\n++++") end end - 执行转换命令:运行Pandoc时指定该过滤器,命令如下:
pandoc input.html -o output.adoc --lua-filter=math-convert.lua - 适配调整:根据你实际的HTML公式结构(比如是否用MathML、是否是内嵌LaTeX文本)修改过滤器的匹配规则,确保能准确提取公式内容。
二、替代Pandoc的转换方案
如果不想依赖Pandoc,还有两种实用方式:
- 自定义Python脚本处理:用BeautifulSoup解析HTML,定位所有数学公式元素,直接生成Asciidoc Stem块:
from bs4 import BeautifulSoup # 读取HTML文件 with open("input.html", "r", encoding="utf-8") as f: soup = BeautifulSoup(f.read(), "html.parser") adoc_content = "" # 假设公式在class为"math-equation"的容器中 for elem in soup.find_all(class_="math-equation"): formula_text = elem.get_text(strip=True) # 生成latexmath Stem块 adoc_content += "[latexmath]\n++++\n" + formula_text + "\n++++\n\n" # 写入Asciidoc文件 with open("output.adoc", "w", encoding="utf-8") as f: f.write(adoc_content) - 使用Asciidoctor生态工具:借助
asciidoctor-math插件的辅助转换能力,先将HTML中的公式导出为纯LaTeX/AsciiMath文本,再手动或半自动嵌入到Asciidoc的Stem块中,适合公式数量不多的场景。
三、验证转换效果
转换完成后,用Asciidoctor渲染文件验证结果:
asciidoctor output.adoc -o preview.html
打开生成的preview.html,确认公式是否正常渲染,确保Stem块的格式完全符合Asciidoc规范(比如标记[latexmath]或[asciimath]与++++包裹的公式内容配对正确)。
内容的提问来源于stack exchange,提问作者Alexey Bogdanov
相关产品推荐
相关产品推荐

