使用BeautifulSoup从同class的<div>提取邮箱和网站遇问题
解决方法
1. 提取class为"IdiaP Me sNsFa"元素中的链接
直接访问contact["href"]报错,是因为这些元素本身不是<a>标签,href属性在它们的子标签里。需要先定位到子级的<a>标签再提取属性:
contacts_container = s_restaurant.find_all(class_="IdiaP Me sNsFa") for contact in contacts_container: a_tag = contact.find('a') if a_tag and 'href' in a_tag.attrs: print(a_tag["href"])
- 用
find('a')定位元素内部的链接标签 - 增加判断逻辑,避免找不到标签或无
href属性时触发报错
2. 提取class为"YnKZo Ci Wc _S C FPPgD"元素中的网站链接
打印contact.a拿不到href,大概率是<a>标签不在直接子级,或需要更精准的过滤:
target_element = s_restaurant.find(class_="YnKZo Ci Wc _S C FPPgD") # 直接筛选带有href属性的a标签 a_tag = target_element.find('a', href=True) if a_tag: print(a_tag["href"])
find('a', href=True)可以直接过滤出包含href属性的有效链接标签- 如果链接是页面动态渲染生成的,需要改用Selenium这类工具先渲染页面再提取
内容的提问来源于stack exchange,提问作者peetman
相关产品推荐
相关产品推荐

