使用XPath爬取Scamadviser数据仅返回Error的问题排查求助
问题修复方案
核心错误原因
- 语法错误:driver初始化行多了多余右括号,
ic变量所在行缩进不符合Python规范,直接触发语法异常 - 滚动操作逻辑错误:
execute_script执行滚动命令的返回值为None,直接在后面调用.get_attribute()属于对空对象操作,直接报错 - XPath定位完全错误:
- 信任分元素的筛选条件
contains(text(),'icon')不符合实际页面元素结构,根本找不到对应元素 - 国家、ISP字段定位的是标签文本本身,不是对应的取值节点,拿到的内容也不符合需求
- 信任分元素的筛选条件
- 等待逻辑无效:仅实例化了
WebDriverWait对象,没有调用until方法等待元素加载完成,页面未渲染完成就查找元素必然失败
修正后可运行代码
import pandas as pd from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC def scam(): chrome_options = webdriver.ChromeOptions() # 可选添加反爬配置 chrome_options.add_argument('--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36') chrome_options.add_argument('--headless=new') # 可选无头模式 trust=[] country = [] isp_country = [] urls=['stackoverflow.com','github.com'] # 修正多余右括号,新版本selenium用options参数替换chrome_options driver=webdriver.Chrome(options=chrome_options) wait = WebDriverWait(driver, 30) for x in urls: driver.get(f'https://www.scamadviser.com/check-website/{x}') try: # 等待信任分元素加载,滚动到可见区域再取值 trust_ele = wait.until(EC.presence_of_element_located((By.XPATH, "//div[contains(@class,'trust__overlay')]"))) driver.execute_script("arguments[0].scrollIntoView();", trust_ele) t = trust_ele.get_attribute('innerText').strip() trust.append(t) # 定位国家标签对应的取值节点 country_ele = wait.until(EC.presence_of_element_located((By.XPATH, "//div[contains(@class,'block__col') and text()='Country']/following-sibling::div[1]"))) driver.execute_script("arguments[0].scrollIntoView();", country_ele) c = country_ele.get_attribute('innerText').strip() country.append(c) # 定位ISP标签对应的取值节点 isp_ele = wait.until(EC.presence_of_element_located((By.XPATH, "//div[contains(@class,'block__col') and text()='ISP']/following-sibling::div[1]"))) ic = isp_ele.get_attribute('innerText').strip() isp_country.append(ic) except Exception as e: # 可选打印异常信息方便排查:print(f"查询{x}出错:{str(e)}") trust.append("Error") country.append("Error") isp_country.append("Error") # 生成结果DataFrame res_dict = {'URL': urls, 'Trust':trust, 'Country': country, 'ISP': isp_country} df = pd.DataFrame(res_dict) driver.quit() return df # 调用测试 print(scam())
额外注意事项
- 若触发网站反爬机制,可添加代理、调整请求间隔、禁用无头模式后再测试
- Selenium 4.0+版本已废弃
find_element_by_xpath方法,统一使用find_element(By.XPATH, 路径)写法
内容的提问来源于stack exchange,提问作者LdM
相关产品推荐
相关产品推荐

