You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium通过XPath提取Scamadviser网站信息无数据返回问题求助

问题修复方案

核心错误点

  • Xpath匹配逻辑错误:Country文本位于div.block__col下的<b>子标签内,原有匹配父节点text()的逻辑无法命中元素;且你需要提取的是标签对应的字段值,不是标签本身的文本
  • JS滚动操作无返回值,链式调用get_attribute直接触发异常
  • 代码存在语法错误:ChromeDriver初始化多写了右括号、ISP提取行缩进错误
  • 未等待目标元素加载完成就执行查找,页面未渲染完成时找不到元素

修正后可运行代码

首先提前导入依赖:

import pandas as pd
from selenium import webdriver
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By

修正后的函数如下,替换你的chromedriver路径为本地实际路径即可运行:

def scam(df):
    chrome_options = webdriver.ChromeOptions()
    # 不需要弹出浏览器可放开下一行注释开启无头模式
    # chrome_options.add_argument('--headless')
    trust = []
    country = [] 
    isp_country = [] 
        
    query = df['URL'].unique().tolist() 
    # 修复初始化多括号语法错误
    driver = webdriver.Chrome('你的chromedriver路径', chrome_options=chrome_options)
    wait = WebDriverWait(driver, 15)
    
    for x in query:
        driver.get(f'https://www.scamadviser.com/check-website/{x}')
        try:
            # 提取数值型信任分
            ts_ele = wait.until(EC.presence_of_element_located(
                (By.XPATH, "//span[contains(@class,'trustscore__score')]")
            ))
            trust.append(int(ts_ele.text.strip()))
            
            # 提取国家:先定位Country标签,再取相邻兄弟节点的字段值
            country_label = wait.until(EC.presence_of_element_located(
                (By.XPATH, "//div[contains(@class,'block__col')]/b[text()='Country']/parent::div/following-sibling::div[1]")
            ))
            country.append(country_label.text.strip())
            
            # 提取ISP信息
            isp_label = wait.until(EC.presence_of_element_located(
                (By.XPATH, "//div[contains(@class,'block__col')]/b[contains(text(),'ISP')]/parent::div/following-sibling::div[1]")
            ))
            isp_country.append(isp_label.text.strip())
        
        except:
            # 出错时填充默认值
            trust.append(0)
            country.append("Error")
            isp_country.append("Error")
            
    # 生成结果表
    res_df = pd.DataFrame({
        'URL': query, 
        'Trustscore': trust, 
        'Country': country, 
        'ISP': isp_country
    })

    driver.quit()
    return res_df

测试验证

使用你提供的测试地址验证效果:

test_df = pd.DataFrame({'URL': ['stackoverflow.com', 'github.com']})
result = scam(test_df)
print(result)

正常运行会返回两个站点的信任分、所属国家、ISP信息,不会再全返回Error。

内容的提问来源于stack exchange,提问作者LdM

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 22:45:03