You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python提取LaTeX文件方程遇阻,寻求可行解决方案

问题描述

我想用Python从LaTeX文件中提取方程,尝试了以下正则表达式代码:

import re
with open('file.tex',encoding='utf-8') as f: data = f.read()
pattern= r'\\begin\{equation*\}(.*?)\\end\{equation*\}'
re.findall(pattern, data, re.S) 

但输出始终为空(我也尝试过re.DOTALL,同样无效)。我还试用了tex2py库,但目前只能提取文本,无法获取方程。最终我希望将提取到的方程通过st.markdown/st.latex展示在Streamlit前端,请问有人有解决思路吗?

解决思路
  • 修正正则表达式:你当前的正则写法存在错误,equation*会被解析为匹配equatio加上任意数量的n,而非匹配equation或equation*环境。正确的正则应该是:

    pattern = r'\\begin\{equation\*?\}(.*?)\\end\{equation\*?\}'
    equations = re.findall(pattern, data, re.DOTALL)
    

    这里\*?表示匹配0个或1个*,刚好覆盖普通equation和带星号的equation*环境。如果你的LaTeX文件包含嵌套子环境(比如equation内的split),可以进一步调整正则适配,但这个修正足以解决基础的空输出问题。

  • 使用专业LaTeX解析库:正则处理复杂LaTeX结构容易遗漏场景,推荐用专门的解析库,比如latexparser(需先执行pip install latexparser安装),它能精准识别LaTeX的各类环境:

    from latexparser import LatexParser
    
    parser = LatexParser()
    doc = parser.parse_file('file.tex')
    equations = []
    for env in doc.environments:
        if env.name in ('equation', 'equation*'):
            equations.append(env.content.strip())
    

    这种方法比正则更可靠,能处理嵌套环境、注释、标签等复杂情况。

  • Streamlit展示方程:提取到方程内容后,直接用st.latex()即可完美渲染:

    import streamlit as st
    
    for eq in equations:
        st.latex(eq)
    

    如果想用st.markdown(),可以给方程加上块级公式包裹符$$:

    st.markdown(f"$$ {eq} $$")
    

内容的提问来源于stack exchange,提问作者Danny

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 17:23:20