You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup提取并输出含LaTeX公式的原始字符串

问题解决方法

你的代码里两次错误使用了repr()函数,导致HTML字符串被额外转义,出现了\x07这种乱码字符(原字符串里的\a被转成了ASCII控制字符)。

错误点分析

  • 初始化BeautifulSoup时,你传入了repr(a),但a本身就是合法的HTML字符串,repr()会把字符串的原始表示(包括转义符)传递给解析器,破坏了原有的HTML结构。
  • 找到div.statement后,又用bs(repr(...))重新解析,这属于重复且错误的操作,进一步加剧了转义问题。

修正后的代码

a = """<div class="theorem" id="theorem-AA" acro="AA" titletext="Adjoint of an Adjoint"> <h5 class="theorem"> <span class="type">Theorem </span><span class="acro">AA</span><span class="titletext"> Adjoint of an Adjoint</span> </h5> <div class="statement"><p>Suppose that $A$ is a matrix.  Then $\adjoint{\left(\adjoint{A}\right)}=A$.</p></div> <div class="proof"><a knowl="./knowls/proof.AA.knowl">Proof</a></div> </div><div class="context"><a href="http://linear.pugetsound.edu/html/section-MO.html#theorem-AA" class="context" title="Section MO">(in context)</a></div> """

from bs4 import BeautifulSoup as bs

# 直接用原始字符串初始化soup,不要用repr()
soup = bs(a, features='lxml')

# 找到第一个div.statement,直接从中提取p标签文本
statement_div = soup.find("div", {"class": "statement"})
statement = statement_div.find('p').text

print(statement)

说明

  • 去掉所有不必要的repr()调用,让BeautifulSoup直接解析原始HTML字符串。
  • 使用find()代替find_all()[0]更简洁,直接获取第一个匹配的元素。

运行修正后的代码,就能得到你想要的正确输出:

Suppose that $A$ is a matrix.  Then $\adjoint{\left(\adjoint{A}\right)}=A$.

内容的提问来源于stack exchange,提问作者Vinod

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 22:12:26