将adoc转换为Markdown时如何保留LaTeX格式数学公式?
Asciidoc转Markdown时保留LaTeX公式的解决方案
方法1:直接用Asciidoctor生成Markdown(跳过DocBook中转)
Asciidoctor原生支持直接输出Markdown格式,无需经过DocBook中转,能精准识别latexmath:[$some_equation_here$]语法并转换为Markdown兼容的$some_equation_here$格式,还能减少后续正则修正的工作量。
使用命令:
# 生成标准Markdown asciidoctor -b markdown -o <outfile> <infile> # 若需严格遵循Markdown规范,指定strict后端 asciidoctor -b markdown_strict -o <outfile> <infile>
方法2:用XSLT预处理DocBook XML(适配原有中转流程)
如果必须保留原有的Asciidoc→DocBook→Pandoc流程,可以通过XSLT脚本提取DocBook中<inlineequation>块的公式内容,再交给Pandoc处理。
- 创建
extract-math.xsl脚本:
<xsl:stylesheet version="1.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform"> <!-- 保留文档其他内容不变 --> <xsl:template match="@*|node()"> <xsl:copy> <xsl:apply-templates select="@*|node()"/> </xsl:copy> </xsl:template> <!-- 提取行内公式的CDATA内容 --> <xsl:template match="inlineequation"> <xsl:value-of select="alt/text()"/> </xsl:template> <!-- 提取多行公式的CDATA内容(针对equation块) --> <xsl:template match="equation"> <xsl:value-of select="alt/text()"/> </xsl:template> </xsl:stylesheet>
- 修改转换命令链:
asciidoc -b docbook -o temp.xml <infile> xsltproc extract-math.xsl temp.xml > temp-processed.xml pandoc -f docbook -t markdown_strict --atx-headers --mathjax temp-processed.xml -o <outfile>
方法3:优化正则处理逻辑(适配现有正则方案)
如果坚持使用正则修正,需调整匹配逻辑,避免双重转义和多余标签问题,注意要在Asciidoc转DocBook之前处理原始adoc文件:
import re def fix_latexmath(content): # 匹配latexmath语法,兼容带/不带美元符号的情况 content = re.sub(r'latexmath:\[(.*?)\]', r'$\1$', content, flags=re.DOTALL) # 移除Asciidoc自动生成的多余<sup>标签 content = re.sub(r'</?sup>', '', content) # 修复双重转义的反斜杠 content = re.sub(r'\\\\', r'\\', content) return content # 使用示例:读取原始adoc文件,处理后写入临时文件 with open('input.adoc', 'r', encoding='utf-8') as f: raw_content = f.read() processed_content = fix_latexmath(raw_content) with open('temp-processed.adoc', 'w', encoding='utf-8') as f: f.write(processed_content)
之后再执行原有的DocBook转换和Pandoc命令即可。
内容的提问来源于stack exchange,提问作者1_Esk
相关产品推荐
相关产品推荐

