You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python和BeautifulSoup提取HTML中的指定标题内容?

你的代码已经定位到了所有class为title的<a>标签,要提取其中的标题内容,分以下几种场景处理:

1. 提取标签内的文本内容(即<a>与</a>之间的"Blah")

使用Tag对象的get_text()方法(推荐,可自动去除前后空白)或.text属性:

修改原代码的打印部分:

titles = soup.select('a.title')
# 循环提取每个标签的文本
for title in titles:
    # strip=True 去除文本前后的空白字符
    print(title.get_text(strip=True))

# 或者用列表推导式一次性获取所有文本
title_texts = [title.get_text(strip=True) for title in titles]
print(title_texts)

2. 提取title属性中的内容(即title="Blah"里的"Blah")

通过get()方法获取属性值(比直接用['title']更安全,避免标签无该属性时报错):

titles = soup.select('a.title')
title_attrs = [title.get('title') for title in titles]
print(title_attrs)

3. 提取title-auto-hide属性中的内容

同样用get()方法:

titles = soup.select('a.title')
title_auto_hide_vals = [title.get('title-auto-hide') for title in titles]
print(title_auto_hide_vals)

选择哪种方式取决于你实际需要的是哪部分"Blah"内容。

内容的提问来源于stack exchange,提问作者Magicskid

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 03:15:38