You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

BeautifulSoup抓取结果如何去除HTML标签仅获取目标纯文本

Python BeautifulSoup 提取纯文本解决方案

你当前代码返回带HTML标签的结果,原因是find_all()返回的是所有匹配标签对象组成的列表,直接打印该列表会输出标签的完整HTML结构,按以下方式修改即可提取纯文本:

最优修改方案

如果确认页面仅存在1个符合条件的h1标签,直接用find()替代find_all(),再调用标签的.text属性提取文本:

from bs4 import BeautifulSoup
import requests

url = 'https://www.mediacorp.sg/en/your-mediacorp/our-artistes/tca/male-artistes/ayden-sng-12357686'

artiste_name = 'celeb-name'

page = requests.get(url)

soup = BeautifulSoup(page.text, 'lxml')

# 修改点:用find获取单个标签,调用text属性取纯文本
txt = soup.find('h1', attrs={'class':artiste_name}).text

print(txt)

兼容find_all的方案

如果需要匹配多个结果的场景,先按索引取出对应标签后再提取文本即可:

# 取列表第一个匹配结果,再提取文本
txt = soup.find_all('h1', attrs={'class':artiste_name})[0].text

内容的提问来源于stack exchange,提问作者munchies

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 18:54:06