You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用BeautifulSoup移除<p>元素中的所有<span>标签?

使用BeautifulSoup移除

中的标签

方法1:删除标签及其内容(用decompose())

如果需要彻底移除<span>标签以及里面的所有文本,调用decompose()方法即可删除目标标签及其子内容。

示例代码:

from bs4 import BeautifulSoup

# 假设你的HTML内容存储在html变量中
html = '''
<div class="myclass">
  <p>
    text 1 to keep<span>text 1 to remove</span>and keep this too.
  </p>
  <p>
    text 2 to keep<span>text 2 to remove</span>and keep this too.
  </p>
</div>
'''

soup = BeautifulSoup(html, 'html.parser')

# 定位所有.myclass下<p>标签内的<span>,逐个删除
for span in soup.select('.myclass p span'):
    span.decompose()

# 提取所有<p>的文本内容
text = ""
for tag in soup.find_all(class_="myclass"):
    for p in tag.find_all('p'):
        text += p.get_text(strip=True) + "\n"  # strip=True去除多余空格,换行分隔不同段落

print(text)

运行结果:

text 1 to keepand keep this too.
text 2 to keepand keep this too.

方法2:仅移除标签,保留其内容(用unwrap())

如果只是想去掉<span>标签外壳,保留里面的文本内容,使用unwrap()方法将标签剥离,内容会留在原位置。

示例代码:

from bs4 import BeautifulSoup

html = '''
<div class="myclass">
  <p>
    text 1 to keep<span>text 1 to remove</span>and keep this too.
  </p>
  <p>
    text 2 to keep<span>text 2 to remove</span>and keep this too.
  </p>
</div>
'''

soup = BeautifulSoup(html, 'html.parser')

for span in soup.select('.myclass p span'):
    span.unwrap()

text = ""
for tag in soup.find_all(class_="myclass"):
    for p in tag.find_all('p'):
        text += p.get_text(strip=True) + "\n"

print(text)

运行结果:

text 1 to keeptext 1 to removeand keep this too.
text 2 to keeptext 2 to removeand keep this too.

补充说明

  • 原代码中tag.p.text仅会提取第一个<p>标签的文本,改为遍历所有<p>标签才能获取全部内容。
  • 使用select('.myclass p span')通过CSS选择器定位,比逐层查找更简洁精准,直接锁定目标<span>标签。

内容的提问来源于stack exchange,提问作者MarcoS

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 08:35:25