Python网页数据提取故障排查:无法获取网站姓名与邮箱
问题解决:从BHHS代理页面提取邮箱为空的问题
核心问题分析
你的脚本生成空文件主要有以下几个原因:
- 动态页面渲染限制:目标网站的代理列表是通过JavaScript动态加载的,
requests.get只能获取初始静态HTML,无法拿到实际的代理链接元素(.cmp-cta),导致后续没有可遍历的详情页链接。 - 请求头缺失:直接用
requests发起请求会被网站识别为非浏览器请求,返回内容不完整或被拦截。 - 代码语法错误:原脚本中
your textfrom bs4 import BeautifulSoup`这行是无效代码,会导致脚本运行报错。 - 选择器容错性不足:即使进入详情页,邮箱元素的类名可能发生变化,或是邮箱通过动态方式插入,单一选择器无法匹配。
修正后的脚本
import requests from bs4 import BeautifulSoup # 模拟浏览器请求头,避免被反爬拦截 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } # 主页面URL url = "https://www.bhhs.com/agent-search-results" response = requests.get(url, headers=headers) if response.status_code == 200: soup = BeautifulSoup(response.text, 'html.parser') enlaces = soup.find_all('a', class_='cmp-cta') # 处理动态加载场景:若静态页面无链接,提示抓包找API(需自行通过浏览器开发者工具抓包确认API地址) if not enlaces: print("警告:静态页面未找到代理链接,页面为动态加载,请通过浏览器抓包获取代理列表API接口") # 示例:替换为实际抓包得到的API地址 # api_url = "https://www.bhhs.com/api/agent-search?page=1&size=20" # api_res = requests.get(api_url, headers=headers) # if api_res.status_code == 200: # agents_data = api_res.json() # enlaces = [agent['profileUrl'] for agent in agents_data.get('results', []) if 'profileUrl' in agent] correos = [] for enlace in enlaces: href = enlace.get('href') if href: # 处理相对链接,转为完整URL if not href.startswith('http'): href = f"https://www.bhhs.com{href}" sub_response = requests.get(href, headers=headers) if sub_response.status_code == 200: sub_soup = BeautifulSoup(sub_response.text, 'html.parser') # 优先按原类名查找,失败则尝试抓取mailto链接 correo_elemento = sub_soup.find('div', class_='cmp-agent-details__mail text-lowercase') if correo_elemento: correo = correo_elemento.text.strip() correos.append(correo) print(f"找到邮箱:{correo}") else: mailto_link = sub_soup.find('a', href=lambda x: x and x.startswith('mailto:')) if mailto_link: correo = mailto_link.get('href').replace('mailto:', '').strip() correos.append(correo) print(f"找到邮箱:{correo}") else: print(f"访问详情页失败 {href},状态码:{sub_response.status_code}") # 保存结果到文件 with open('correos.txt', 'w', encoding='utf-8') as file: for correo in correos: file.write(f"{correo}\n") print(f"提取完成,共找到{len(correos)}个邮箱,已保存到correos.txt") else: print(f"请求主页面失败,状态码:{response.status_code}")
关键优化点说明
- 请求头配置:添加
User-Agent模拟浏览器请求,降低被反爬拦截的概率。 - 动态页面兼容:若静态页面无法获取链接,需通过浏览器开发者工具的「网络」面板抓包,找到网站加载代理列表的API接口,直接请求API获取数据(这是处理动态页面最可靠的方式)。
- 容错性选择器:同时支持按类名和
mailto链接提取邮箱,避免因页面结构变化导致提取失败。 - 相对链接处理:自动将相对路径转为完整URL,避免请求无效地址。
内容的提问来源于stack exchange,提问作者Kalori
相关产品推荐
相关产品推荐

