You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python网页数据提取故障排查:无法获取网站姓名与邮箱

问题解决:从BHHS代理页面提取邮箱为空的问题

核心问题分析

你的脚本生成空文件主要有以下几个原因:

  1. 动态页面渲染限制:目标网站的代理列表是通过JavaScript动态加载的,requests.get只能获取初始静态HTML,无法拿到实际的代理链接元素(.cmp-cta),导致后续没有可遍历的详情页链接。
  2. 请求头缺失:直接用requests发起请求会被网站识别为非浏览器请求,返回内容不完整或被拦截。
  3. 代码语法错误:原脚本中your textfrom bs4 import BeautifulSoup`这行是无效代码,会导致脚本运行报错。
  4. 选择器容错性不足:即使进入详情页,邮箱元素的类名可能发生变化,或是邮箱通过动态方式插入,单一选择器无法匹配。

修正后的脚本

import requests
from bs4 import BeautifulSoup

# 模拟浏览器请求头,避免被反爬拦截
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
}

# 主页面URL
url = "https://www.bhhs.com/agent-search-results"

response = requests.get(url, headers=headers)

if response.status_code == 200:
    soup = BeautifulSoup(response.text, 'html.parser')
    enlaces = soup.find_all('a', class_='cmp-cta')
    
    # 处理动态加载场景:若静态页面无链接,提示抓包找API(需自行通过浏览器开发者工具抓包确认API地址)
    if not enlaces:
        print("警告:静态页面未找到代理链接,页面为动态加载,请通过浏览器抓包获取代理列表API接口")
        # 示例:替换为实际抓包得到的API地址
        # api_url = "https://www.bhhs.com/api/agent-search?page=1&size=20"
        # api_res = requests.get(api_url, headers=headers)
        # if api_res.status_code == 200:
        #     agents_data = api_res.json()
        #     enlaces = [agent['profileUrl'] for agent in agents_data.get('results', []) if 'profileUrl' in agent]
    
    correos = []
    
    for enlace in enlaces:
        href = enlace.get('href')
        if href:
            # 处理相对链接,转为完整URL
            if not href.startswith('http'):
                href = f"https://www.bhhs.com{href}"
            
            sub_response = requests.get(href, headers=headers)
            if sub_response.status_code == 200:
                sub_soup = BeautifulSoup(sub_response.text, 'html.parser')
                # 优先按原类名查找,失败则尝试抓取mailto链接
                correo_elemento = sub_soup.find('div', class_='cmp-agent-details__mail text-lowercase')
                if correo_elemento:
                    correo = correo_elemento.text.strip()
                    correos.append(correo)
                    print(f"找到邮箱:{correo}")
                else:
                    mailto_link = sub_soup.find('a', href=lambda x: x and x.startswith('mailto:'))
                    if mailto_link:
                        correo = mailto_link.get('href').replace('mailto:', '').strip()
                        correos.append(correo)
                        print(f"找到邮箱:{correo}")
            else:
                print(f"访问详情页失败 {href},状态码:{sub_response.status_code}")
    
    # 保存结果到文件
    with open('correos.txt', 'w', encoding='utf-8') as file:
        for correo in correos:
            file.write(f"{correo}\n")
    
    print(f"提取完成,共找到{len(correos)}个邮箱,已保存到correos.txt")
else:
    print(f"请求主页面失败,状态码:{response.status_code}")

关键优化点说明

  • 请求头配置:添加User-Agent模拟浏览器请求,降低被反爬拦截的概率。
  • 动态页面兼容:若静态页面无法获取链接,需通过浏览器开发者工具的「网络」面板抓包,找到网站加载代理列表的API接口,直接请求API获取数据(这是处理动态页面最可靠的方式)。
  • 容错性选择器:同时支持按类名和mailto链接提取邮箱,避免因页面结构变化导致提取失败。
  • 相对链接处理:自动将相对路径转为完整URL,避免请求无效地址。

内容的提问来源于stack exchange,提问作者Kalori

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 20:25:17