You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup和lxml解析时为何获取到多余文本?

问题分析与解决

问题详情

HTML结构

<div class="body">

       <div class="pull_right date details" title="21.11.2024 20:17:23 UTC+07:00">
20:17
       </div>

       <div class="from_name">
Cheki_FNS 
       </div>

       <div class="text">
Cash receipt received: from <strong>Komandor trading network</strong> (LLC &quot;TS KOMANDOR&quot;)
       </div>

      </div>

     </div>

所用代码

with open("messages.html", "r", encoding="utf-8") as file:
    html_content = file.read()

soup = BeautifulSoup(html_content, "xml")

div_text = soup.find("div", class_="text")


if div_text:
    print(div_text.get_text())
else:
    print("Error, class text not find")

结果对比

  • 预期提取结果:
"A cash receipt has been received: from the Komandor Trading Network (TS Komandor LLC)"
  • 实际得到结果:
"20:17  Cheki_FNS A cash receipt has been received: from the Komandor Trading Network (TS Komandor LLC)". 

用户疑问:是否是部分文本超出div对象范围导致的该问题?


解答

不是文本超出div范围的问题,核心原因是错误使用了xml解析器处理HTML内容:
XML解析器对HTML标签结构兼容性差,会保留并合并HTML里的换行、空格等空白文本节点,甚至错误识别DOM结构,最终导致get_text()返回了多个div的文本内容。

修复方案

将解析器替换为专门处理HTML的解析器,比如内置的html.parser,或者第三方的lxml(需额外安装)。同时可以用strip=True参数清理文本中的多余空白。

修改后的代码:

with open("messages.html", "r", encoding="utf-8") as file:
    html_content = file.read()

# 改用HTML专用解析器
soup = BeautifulSoup(html_content, "html.parser")

div_text = soup.find("div", class_="text")

if div_text:
    # strip=True 去除首尾空白,同时合并内部多余空格
    print(div_text.get_text(strip=True))
else:
    print("Error, class text not find")

运行后即可得到仅来自class="text"的div的文本内容,符合预期。


内容的提问来源于stack exchange,提问作者ARBYZIK

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.15 23:13:14