You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用XPath爬取Scamadviser数据仅返回Error的问题排查求助

问题修复方案

核心错误原因

  • 语法错误:driver初始化行多了多余右括号,ic变量所在行缩进不符合Python规范,直接触发语法异常
  • 滚动操作逻辑错误:execute_script执行滚动命令的返回值为None,直接在后面调用.get_attribute()属于对空对象操作,直接报错
  • XPath定位完全错误:
    • 信任分元素的筛选条件contains(text(),'icon')不符合实际页面元素结构,根本找不到对应元素
    • 国家、ISP字段定位的是标签文本本身,不是对应的取值节点,拿到的内容也不符合需求
  • 等待逻辑无效:仅实例化了WebDriverWait对象,没有调用until方法等待元素加载完成,页面未渲染完成就查找元素必然失败

修正后可运行代码

import pandas as pd
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

def scam():
    chrome_options = webdriver.ChromeOptions()
    # 可选添加反爬配置
    chrome_options.add_argument('--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36')
    chrome_options.add_argument('--headless=new') # 可选无头模式

    trust=[]
    country = [] 
    isp_country = [] 
    urls=['stackoverflow.com','github.com']
    # 修正多余右括号,新版本selenium用options参数替换chrome_options
    driver=webdriver.Chrome(options=chrome_options)
    wait = WebDriverWait(driver, 30)
    
    for x in urls:
        driver.get(f'https://www.scamadviser.com/check-website/{x}')
        try: 
            # 等待信任分元素加载,滚动到可见区域再取值
            trust_ele = wait.until(EC.presence_of_element_located((By.XPATH, "//div[contains(@class,'trust__overlay')]")))
            driver.execute_script("arguments[0].scrollIntoView();", trust_ele)
            t = trust_ele.get_attribute('innerText').strip()
            trust.append(t)  

            # 定位国家标签对应的取值节点
            country_ele = wait.until(EC.presence_of_element_located((By.XPATH, "//div[contains(@class,'block__col') and text()='Country']/following-sibling::div[1]")))
            driver.execute_script("arguments[0].scrollIntoView();", country_ele)
            c = country_ele.get_attribute('innerText').strip()
            country.append(c)  

            # 定位ISP标签对应的取值节点
            isp_ele = wait.until(EC.presence_of_element_located((By.XPATH, "//div[contains(@class,'block__col') and text()='ISP']/following-sibling::div[1]")))
            ic = isp_ele.get_attribute('innerText').strip()
            isp_country.append(ic)
        
        except Exception as e:
            # 可选打印异常信息方便排查:print(f"查询{x}出错:{str(e)}")
            trust.append("Error")
            country.append("Error")
            isp_country.append("Error")
            
    # 生成结果DataFrame
    res_dict = {'URL': urls, 'Trust':trust, 'Country': country, 'ISP': isp_country} 
    df = pd.DataFrame(res_dict)
    driver.quit()
    return df

# 调用测试
print(scam())

额外注意事项

  • 若触发网站反爬机制,可添加代理、调整请求间隔、禁用无头模式后再测试
  • Selenium 4.0+版本已废弃find_element_by_xpath方法,统一使用find_element(By.XPATH, 路径)写法

内容的提问来源于stack exchange,提问作者LdM

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 07:45:05