如何用BeautifulSoup提取h2标签内a标签的title属性?
解决BeautifulSoup获取h2内a标签title属性的问题
嘿,我来帮你搞定这个抓取问题!你之前的代码没能拿到目标属性,核心原因是找错了属性所在的标签——title是在<h2>里面的<a>标签上,而不是<h2>本身,所以直接筛选带title的h2肯定行不通~
下面给你两种可行的解决方案,代码都经过验证:
方法1:使用CSS选择器直接定位目标标签
CSS选择器可以非常直观地定位到<h2>下的所有<a>标签,代码简洁高效:
from bs4 import BeautifulSoup import urllib2 url = "http://www.example.it" page = urllib2.urlopen(url) soup = BeautifulSoup(page, 'html.parser') # 用CSS选择器选中所有h2标签下的a标签 target_links = soup.select('h2 a') # 遍历获取每个a标签的title属性 for link in target_links: # 使用get方法避免属性不存在时抛出错误 title = link.get('title') if title: print(title)
方法2:先找h2再嵌套查找内部的a标签
如果你习惯用BeautifulSoup的find系列方法,也可以先获取所有h2标签,再逐个查找里面的a标签:
from bs4 import BeautifulSoup import urllib2 url = "http://www.example.it" page = urllib2.urlopen(url) soup = BeautifulSoup(page, 'html.parser') # 获取所有h2标签 h2_tags = soup.findAll('h2') for h2 in h2_tags: # 查找当前h2下的第一个a标签(如果有多个可以用findAll) a_tag = h2.find('a') if a_tag: title = a_tag.get('title') if title: print(title)
为什么之前的方法失效?
你尝试的findAll('h2', attrs={'title'})逻辑是查找自身带有title属性的h2标签,但你的目标title属性是在h2的子元素a标签上,所以这个筛选条件完全匹配不到内容,自然拿不到结果啦。
另外提醒你:用tag.get('属性名')比直接tag['属性名']更安全,因为如果某个a标签没有title属性,直接取值会抛出KeyError,而get方法会返回None,让程序更健壮。
内容的提问来源于stack exchange,提问作者heisen
相关产品推荐
相关产品推荐

