You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Selenium提取邮箱异常:部分网站无法捕获邮箱问题求助

Selenium提取网站邮箱不稳定问题的解决方法

问题根源分析

你的代码仅通过//a[contains(@href, 'mailto:')]定位邮箱,只能捕获封装在邮件链接里的地址。但restauranteelpicaporte.es的邮箱是以纯文本形式展示的,没有套在mailto链接中,所以当前逻辑抓不到。另外,固定的time.sleep(2)依赖网络环境,页面没加载完就执行查找会导致不稳定。

改进后的代码

下面的代码同时支持两种捕获方式:先找mailto链接,找不到就用正则提取页面中符合邮箱格式的文本,同时用显式等待替代固定sleep提升稳定性。

import re
import time
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.chrome.options import Options

# 浏览器配置优化
options = Options()
options.add_argument('--start-maximized')
options.add_argument('--disable-extensions')
options.add_argument('--disable-blink-features=AutomationControlled')  # 规避反爬检测
options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36")

driver_path = 'C:\\Users\\chromedriver_win32\\chromedriver.exe'
driver = webdriver.Chrome(service=webdriver.chrome.service.Service(driver_path), options=options)
wait = WebDriverWait(driver, 10)  # 延长等待时间,适配不同加载速度

def get_email(url):
    driver.execute_script("window.open('');")
    driver.switch_to.window(driver.window_handles[-1])
    try:
        driver.get(url)
        # 等待页面主体加载完成
        wait.until(EC.presence_of_element_located((By.TAG_NAME, 'body')))
        
        # 优先捕获mailto链接
        try:
            email_element = wait.until(EC.presence_of_element_located((By.XPATH, "//a[contains(@href, 'mailto:')]")))
            email_addr = email_element.get_attribute("href").replace("mailto:", "").strip()
            return email_addr
        except:
            # 正则提取页面文本中的邮箱
            # 滚动到页面底部,确保动态加载的内容都显示
            driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
            time.sleep(1)
            
            page_source = driver.page_source
            # 匹配标准邮箱格式
            email_pattern = r'[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}'
            matches = re.findall(email_pattern, page_source)
            
            if matches:
                # 去重并过滤无效匹配
                unique_emails = list(set(matches))
                valid_emails = [email for email in unique_emails if '@' in email and len(email.split('@')[-1]) >= 3]
                return valid_emails[0] if valid_emails else ""
            return ""
    finally:
        # 无论成功失败都关闭标签页切回原窗口
        driver.close()
        driver.switch_to.window(driver.window_handles[0])

# 测试代码
website1 = 'restauranteelpicaporte.es'  # INFO@RESTAURANTEELPICAPORTE.ES
website2 = 'restaurante-laparra.com'    # laparramadrid@gmail.com

driver.get('http://www.google.es/maps/')
time.sleep(2)

email1 = get_email('https://' + website1)
email2 = get_email('https://' + website2)

print("email1:", email1)
print("email2:", email2)

额外优化建议

  • 处理iframe中的邮箱:如果目标邮箱在iframe内,需要先切换到iframe再执行查找:
    iframe = wait.until(EC.presence_of_element_located((By.XPATH, "//iframe[@id='target-iframe']")))
    driver.switch_to.frame(iframe)
    # 执行查找操作...
    driver.switch_to.default_content()  # 切回主文档
    
  • 多页面并行处理:如果需要批量抓取,可考虑使用多线程或异步加载提升效率,但注意控制请求频率避免被封。
  • 更精准的正则:如果目标网站邮箱有特定格式,可调整正则表达式减少无效匹配。

内容的提问来源于stack exchange,提问作者Jaime Martinez

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 15:22:50