Scrapy+Selenium自动化登录报错:WebDriver无find_element_by_xpath属性
解决Scrapy-Selenium中WebDriver无find_element_by_xpath属性的问题
问题背景
使用Scrapy结合Selenium实现站点自动登录,已配置settings.py并编写爬虫代码,浏览器可正常弹出并加载页面,但执行元素输入操作时触发报错:
File "D:\SCRAPPING\tpad_reports\tpad_reports\spiders\tsc_reports.py", line 23, in parse
tsc_no=driver.find_element_by_xpath('//input[@type="text"]')
AttributeError: 'WebDriver' object has no attribute 'find_element_by_xpath'
原因分析
该错误是由于Selenium 4.x及以上版本移除了find_element_by_xpath、find_element_by_id这类旧版元素定位API,统一改为使用find_element()方法配合By类指定定位策略的新写法。代码中仍在调用已被废弃的旧API,因此触发属性不存在的错误。
解决方案
- 导入Selenium的
By类,用于指定定位策略 - 替换所有旧版定位方法为
driver.find_element(By.定位方式, 定位表达式)的格式 - 推荐使用显式等待替代
time.sleep(),提升代码稳定性(可选但建议)
修正后的完整爬虫代码
import scrapy from scrapy_selenium import SeleniumRequest from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC class TscReportsSpider(scrapy.Spider): name = 'tsc_reports' allowed_domains = ['tpad2.tsc.go.ke'] start_urls = ['https://tpad2.tsc.go.ke/auth/login'] def start_requests(self): for url in self.start_urls: # 用显式等待替代自定义wait函数,等待页面元素加载完成 yield SeleniumRequest( url=url, wait_time=10, wait_until=EC.presence_of_element_located((By.XPATH, '//input[@type="text"]')), callback=self.parse ) def parse(self, response): driver = response.request.meta['driver'] # 用新API定位元素并等待可交互 tsc_no = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.XPATH, '//input[@type="text"]')) ) tsc_no.send_keys(620127) id_no = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.XPATH, '//input[@type="number"]')) ) id_no.send_keys(27376193) password = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.XPATH, '//input[@type="password"]')) ) password.send_keys(620127) login_btn = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.XPATH, '//button[@class="btn btn-md"]')) ) login_btn.click() # 用click()替代submit(),逻辑更直观 # 可选:等待登录后的页面加载,替换为实际页面的特征(如标题关键词) WebDriverWait(driver, 10).until( EC.title_contains("TPAD") )
关键修改点说明
- 移除了自定义
wait函数和time.sleep(),改用WebDriverWait配合expected_conditions实现显式等待,确保元素加载完成后再操作,避免因页面加载慢导致的元素找不到问题 - 将所有旧版
find_element_by_xpath()调用替换为新版API,既兼容Selenium 4.x+,又提升代码稳定性 - 用
click()方法点击登录按钮,替代submit(),逻辑更清晰直观
验证步骤
- 通过
pip show selenium确认Selenium版本为4.x及以上 - 替换修正后的爬虫代码
- 在命令行执行
scrapy runspider tsc_reports.py - 观察浏览器是否正常完成输入并登录操作
内容的提问来源于stack exchange,提问作者JuliusFx
相关产品推荐
相关产品推荐

