You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium爬虫提取邮箱网站返回None的XPath修复方案

Selenium爬取律师平台联系方式字段返回None修复

问题场景

运行以下Selenium脚本爬取荷兰律师查询平台的律师公开信息时,进入律师个人详情页后,现有XPath规则提取邮箱、个人官网字段均返回None值,无法拿到目标数据。

import time
from selenium import webdriver
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.common.by import By
from selenium.webdriver.support.wait import WebDriverWait
from webdriver_manager.chrome import ChromeDriverManager

options = webdriver.ChromeOptions()
options.add_argument("--no-sandbox")
options.add_argument("--disable-gpu")
options.add_argument("--window-size=1920x1080")
options.add_argument("--disable-extensions")

chrome_driver = webdriver.Chrome(
    service=Service(ChromeDriverManager().install()),
    options=options
)


def supplyvan_scraper():
    with chrome_driver as driver:
        driver.implicitly_wait(15)
        URL = 'https://zoekeenadvocaat.advocatenorde.nl/zoeken?q=&type=advocaten&limiet=10&sortering=afstand&filters[rechtsgebieden]=[]&filters[specialisatie]=0&filters[toevoegingen]=0&locatie[adres]=Holland&locatie[geo][lat]=52.132633&locatie[geo][lng]=5.291266&locatie[straal]=56&locatie[hash]=67eb2b8d0aab60ec69666532ff9527c9&weergave=lijst&pagina=1'
        driver.get(URL)
        time.sleep(3)

        page_links = [element.get_attribute('href') for element in
                      driver.find_elements(By.XPATH, "//span[@class='h4 no-margin-bottom']//a")]

        # 遍历所有律师详情页链接
        for link in page_links:
            driver.get(link)
            time.sleep(2)
            try:
                title = driver.find_element(By.CSS_SELECTOR, '.title h3').text
            except:
                pass
            
            
            details=driver.find_elements(By.XPATH,"//section[@class='lawyer-info']")
            for detail in details:
                try:
                    email=detail.find_element(By.XPATH, "//div[@class='row'][3]//div[@class='column small-9']").get_attribute('href')
                except:
                    pass
                try:
                    website=detail.find_element(By.XPATH, "//div[@class='row'][4]//div[@class='column small-9']").get_attribute('href')
                except:
                    pass
                print(title,email,website)
            time.sleep(2)

        time.sleep(2)
        driver.quit()


supplyvan_scraper()

页面结构参考截图:
页面结构截图

错误原因

  • 原有XPath固定取第3、第4个row节点,不同律师的信息条目数量存在差异,固定序号会出现定位偏移
  • 定位目标为外层div节点,div本身不存在href属性,自然返回None
  • 子元素查找的XPath未加.前缀,会从整个页面根节点全局匹配,而非从当前detail节点下查找,容易匹配到无关节点

修复方案

替换详情页内邮箱、官网的提取逻辑,通过字段文本特征定位对应行,再提取行内a标签的链接,核心修改代码如下:

for link in page_links:
    driver.get(link)
    time.sleep(2)
    title = ""
    try:
        title = driver.find_element(By.CSS_SELECTOR, '.title h3').text
    except:
        pass
    
    details=driver.find_elements(By.XPATH,"//section[@class='lawyer-info']")
    for detail in details:
        email = ""
        website = ""
        # 定位邮箱:匹配包含荷兰语“E-mailadres”文本的行,取行内a标签的href
        try:
            email = detail.find_element(By.XPATH, ".//div[contains(@class,'row') and .//*[contains(normalize-space(text()),'E-mailadres')]]//a").get_attribute('href').replace('mailto:', '')
        except:
            pass
        # 定位官网:匹配包含荷兰语“Website”文本的行,取行内a标签的href
        try:
            website = detail.find_element(By.XPATH, ".//div[contains(@class,'row') and .//*[contains(normalize-space(text()),'Website')]]//a").get_attribute('href')
        except:
            pass
        print(title,email,website)
    time.sleep(2)

写法说明

  • XPath开头加.限定查找范围为当前detail节点下,避免全局匹配错误
  • 用normalize-space()处理文本前后的空格、换行,避免因为文本留白导致匹配失败
  • 不依赖固定行序号,通过字段的固定文本特征定位,适配所有律师详情页的结构
  • 邮箱的href属性默认带mailto:前缀,替换后可得到纯邮箱地址
  • 提前给title、email、website设置空字符串默认值,避免字段不存在时触发未定义变量报错

内容的提问来源于stack exchange,提问作者Amen Aziz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 21:31:11