You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用BeautifulSoup忽略元素提取指定HTML结构的文本?

用BeautifulSoup提取指定HTML的纯文本内容

完全可以从你提供的HTML结构中提取出目标文本,具体实现方法如下:

步骤说明

  1. 先确保已安装beautifulsoup4库,未安装的话执行以下命令:
pip install beautifulsoup4
  1. 用BeautifulSoup解析HTML,通过内置的文本提取方法获取并清理内容:

代码实现

方法一:split+join清理空格

from bs4 import BeautifulSoup

# 目标HTML内容
html = """<font>
   <em>X</em>
   and
   <em>Y</em>
</font>"""

# 初始化解析器
soup = BeautifulSoup(html, "html.parser")

# 提取文本并清理多余空白
raw_text = soup.get_text()
clean_text = ' '.join(raw_text.split())
print(f'output = "{clean_text}"')

方法二:利用get_text参数直接处理

BeautifulSoup的get_text方法支持参数配置,能更简洁地完成提取和清理:

from bs4 import BeautifulSoup

html = """<font>
   <em>X</em>
   and
   <em>Y</em>
</font>"""

soup = BeautifulSoup(html, "html.parser")
clean_text = soup.get_text(strip=True, separator=' ')
print(f'output = "{clean_text}"')

原理说明

  • soup.get_text()会递归提取所有标签内的文本内容,包括换行和冗余空格。
  • 方法一中的split()会按任意空白字符分割文本,自动忽略连续的换行、空格;' '.join()用单个空格拼接内容,得到规整的文本。
  • 方法二中的strip=True移除文本首尾的空白,separator=' '指定用空格分隔不同文本块,一步完成提取和清理。

内容的提问来源于stack exchange,提问作者kaan46

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 08:55:25