如何使用BeautifulSoup提取HTML元素的title属性内容?
提取div标签的title属性内容
我尝试从以下HTML代码中提取div标签的title属性内容,但至今未成功。
HTML代码片段:
<div class="pull_right date details" title="22.12.2022 01:49:03 UTC-03:00">
我的代码如下:
from bs4 import BeautifulSoup with open("messages.html") as fp: soup = BeautifulSoup(fp, 'html.parser') results = soup.find_all('div', attrs={'class':'pull_right date details'}) print(results)
当前输出是包含所有符合条件的div元素的列表。
解决方法
你的代码已经成功定位到目标div元素,只需进一步提取它们的title属性即可,以下是两种实现方式:
方式1:遍历元素列表逐个提取
from bs4 import BeautifulSoup with open("messages.html") as fp: soup = BeautifulSoup(fp, 'html.parser') results = soup.find_all('div', attrs={'class':'pull_right date details'}) # 遍历每个元素,提取title属性 for item in results: print(item.get('title'))
方式2:用列表推导式批量获取属性值
如果只需要所有title属性的内容,可以直接用列表推导式一次性提取:
from bs4 import BeautifulSoup with open("messages.html") as fp: soup = BeautifulSoup(fp, 'html.parser') titles = [item.get('title') for item in soup.find_all('div', attrs={'class':'pull_right date details'})] print(titles)
补充说明
find_all()返回的是Tag对象组成的列表,每个Tag对象对应一个匹配的div元素。Tag.get('属性名')是安全的属性获取方式:如果元素不存在该属性,会返回None,不会抛出错误。- 也可以用
Tag['属性名']直接获取属性,但元素无该属性时会触发KeyError,建议优先使用get()方法。
内容的提问来源于stack exchange,提问作者Marco Almeida
相关产品推荐
相关产品推荐

