You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Pandoc将LaTeX转HTML时无法保留代码格式的问题求助

解决Pandoc转换LaTeX到HTML时代码格式丢失的问题

先确认LaTeX代码块的写法

Pandoc对LaTeX代码块的识别依赖于标准环境,先检查你的book.tex里代码块是否用了以下两种常用格式:

1. 使用verbatim环境

\begin{verbatim}
def hello():
    print("Hello World")
\end{verbatim}

2. 使用listings包的lstlisting环境

\usepackage{listings}

\begin{lstlisting}[language=Python]
def hello():
    print("Hello World")
\end{lstlisting}

方案一:无需Lua过滤器,直接用Pandoc内置参数

针对verbatim环境

直接运行带完整文档生成参数的命令:

pandoc -s book.tex -o book.html

-s参数会生成包含<head>等结构的完整HTML文档,Pandoc默认会把verbatim环境转成<pre><code>标签,自动保留换行和缩进。

针对lstlisting环境

需要添加--listings参数让Pandoc识别listings环境,还可以配合语法高亮:

pandoc -s --listings book.tex -o book.html

这条命令会将代码块转成带语法高亮的<pre><code>结构,完全保留原格式。


方案二:修复你的Lua过滤器(如果必须使用)

如果之前的Lua过滤器只生成了<pre>标签但没保留格式,大概率是没正确提取代码内容或处理换行/空格。用下面的过滤器替换transform.lua:

-- 转义HTML特殊字符,避免破坏结构
local function escape_html(s)
  return s:gsub('&', '&amp;'):gsub('<', '&lt;'):gsub('>', '&gt;'):gsub('"', '&quot;'):gsub("'", '&#39;')
end

-- 处理Pandoc识别的CodeBlock元素
function CodeBlock(block)
  return pandoc.RawBlock('html', '<pre><code>' .. escape_html(block.text) .. '</code></pre>')
end

-- 处理LaTeX原生的verbatim环境
function RawBlock(el)
  if el.format == 'latex' and el.text:match('^\\begin{verbatim}') then
    -- 提取verbatim内部的代码内容,去掉环境标签
    local code_content = el.text:gsub('^\\begin{verbatim}\n?', ''):gsub('\n?\\end{verbatim}$', '')
    return pandoc.RawBlock('html', '<pre><code>' .. escape_html(code_content) .. '</code></pre>')
  end
end

然后运行命令:

pandoc -s --lua-filter=transform.lua book.tex -o book.html

这个过滤器会精准提取代码内容,保留所有换行、缩进,同时转义特殊字符避免HTML结构出错。


内容的提问来源于stack exchange,提问作者Manvar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 16:12:40