使用BeautifulSoup可爬取姓名但无法爬取邮箱的技术求助
爬虫邮箱定位问题解决方案
问题背景
我编写的基于BeautifulSoup的Python爬虫可以成功抓取目标页面的姓名数据,但始终无法正确定位并爬取邮箱信息,尝试多种方法均未解决。页面元素结构显示:邮箱位于asta-member_title容器下的<a>标签内,标签的href属性带有mailto:前缀。
原代码如下:
import csv import requests from bs4 import BeautifulSoup url = "https://www.asta.org/membership/directory-search-details?memId=900312276" response = requests.get(url) soup = BeautifulSoup(response.content, "html.parser") data = [] for item in soup.find_all("div", {"class": "asta-member"}): name = item.find("h4", {"class": "asta-member__name"}).text data.append([name]) for item in soup.find_all("div", {"class": "asta-member_title"}): email = item.find("div")[2].text data.append([email]) with open("contacts.csv", "w", newline="") as f: writer = csv.writer(f) writer.writerow(["name", "email"]) writer.writerows(data)
问题分析
- 原代码通过索引
[2]取子元素的方式极不稳定,页面结构稍有变动就会失效,且实际邮箱并不在<div>子元素中,而是在<a>标签里。 - 姓名和邮箱分开遍历存储,导致最终写入CSV时,姓名和邮箱各占一行,不符合预期格式。
修正后的代码
import csv import requests from bs4 import BeautifulSoup url = "https://www.asta.org/membership/directory-search-details?memId=900312276" # 添加请求头模拟浏览器,避免被网站拦截或返回不同内容 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } response = requests.get(url, headers=headers) soup = BeautifulSoup(response.content, "html.parser") data = [] # 找到会员容器 member_container = soup.find("div", {"class": "asta-member"}) if member_container: # 获取姓名 name_elem = member_container.find("h4", {"class": "asta-member__name"}) name = name_elem.text.strip() if name_elem else "未获取到姓名" # 获取邮箱:定位到asta-member_title下href包含mailto的a标签 title_container = member_container.find("div", {"class": "asta-member_title"}) email = "未获取到邮箱" if title_container: email_elem = title_container.find("a", href=lambda x: x and "mailto:" in x) if email_elem: # 提取邮箱地址,去掉mailto:前缀(如果text直接显示邮箱则无需额外处理) email = email_elem.text.strip() # 将姓名和邮箱作为一行存入数据 data.append([name, email]) # 写入CSV文件 with open("contacts.csv", "w", newline="", encoding="utf-8") as f: writer = csv.writer(f) writer.writerow(["name", "email"]) writer.writerows(data)
关键改进点
- 使用属性选择器
href=lambda x: x and "mailto:" in x精准定位邮箱链接,避免依赖元素索引的脆弱性。 - 合并姓名和邮箱的获取逻辑,确保每条数据是包含姓名和邮箱的完整行。
- 添加元素存在性检查,避免因元素缺失导致代码报错。
- 增加请求头
User-Agent,模拟浏览器请求,提升爬取成功率。
内容的提问来源于stack exchange,提问作者eInvincible
相关产品推荐
相关产品推荐

