Selenium中Xpath无法获取正确数量元素的问题求助
解决Dexcheck爬虫获取Realized ROI%数据不全的问题
问题根源分析
- 网站更新后懒加载逻辑变更:原全局PAGE_DOWN滚动无法触发表格内部的全部元素加载,表格大概率是独立的滚动容器
- 依赖位置的Xpath不稳定:
(((//div[@class="py-0.5"]/div/p)[position() mod 3=2])/text())[position() mod 2=1]这种基于位置取模的定位,一旦页面元素结构微调就会失效 - 无头模式渲染差异:手动缩放25%生效但Selenium中无效,是因为无头模式下视口渲染逻辑不同,原缩放设置未正确应用
- 固定sleep不可靠:硬编码等待时间无法适配网络波动或页面加载延迟,导致元素未完全加载就开始解析
具体解决方案
1. 针对表格容器的精准滚动加载
替换原全局滚动逻辑,找到表格的滚动容器,用JavaScript滚动到底部,确保所有行都被加载:
def scroll_to_load(driver): try: # 定位表格的滚动容器(根据实际页面结构调整,优先找带overflow-y属性的div) scroll_container = driver.find_element(By.XPATH, '//div[@class="crypto-pnl-table"]//div[contains(@style, "overflow-y")]') last_height = driver.execute_script("return arguments[0].scrollHeight", scroll_container) while True: # 滚动到容器底部 driver.execute_script("arguments[0].scrollTop = arguments[0].scrollHeight", scroll_container) time.sleep(2) new_height = driver.execute_script("return arguments[0].scrollHeight", scroll_container) if new_height == last_height: break last_height = new_height except Exception as e: print(f"滚动加载失败: {e}")
2. 重构Xpath为语义化定位
不再依赖元素位置,而是根据"Realized ROI %"列的关联元素定位数值,避免结构变动影响:
# 在ScrapeData函数中替换Realized_Profit的Xpath Realized_Profit = response.xpath('//div[contains(text(), "Realized ROI %")]/ancestor::div[contains(@class, "table-header")]/following-sibling::div//div[contains(@class, "text-right") and (contains(@class, "text-green-500") or contains(@class, "text-red-500"))]/text()').getall()
注:需根据实际页面元素的class调整,比如检查数值元素的颜色class是否为text-green-500/red-500,确保定位精准
3. 修复无头模式的渲染问题
在get_driver函数中,正确设置缩放比例和新版无头模式,让渲染更接近正常浏览器:
def get_driver(): options = Options() # 更新UA为最新Chrome版本,避免被识别为旧浏览器 options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36") options.set_capability("pageLoadStrategy", "eager") # 提前加载,提升爬虫效率 options.add_argument("window-size=1920x1080") # 设置足够大的视口 options.add_argument("--headless=new") # 使用新版无头模式,渲染更接近正常浏览器 options.add_argument("--force-device-scale-factor=0.25") # 强制设置25%缩放 prefs = {"profile.managed_default_content_settings.images": 2, "permissions.default.stylesheet": 2} options.add_experimental_option("prefs", prefs) driver = Driver(uc=True) return driver
4. 替换固定sleep为显式等待
使用Selenium的WebDriverWait等待关键元素加载完成,避免硬编码等待时间:
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC def scraper(address, driver): data_combined = {'Wallet Address': address} for x in [30, 7, 1]: url = f'https://dexcheck.ai/app/address-analyzer/{address}?chain=eth&timeframe={x}' driver.get(url) # 等待表格容器加载完成 WebDriverWait(driver, 30).until( EC.presence_of_element_located((By.XPATH, '//div[@class="crypto-pnl-table"]')) ) scroll_to_load(driver) # 等待所有Realized ROI元素加载完成 WebDriverWait(driver, 15).until( EC.presence_of_all_elements_located((By.XPATH, '//div[contains(text(), "Realized ROI %")]/ancestor::div[contains(@class, "table-header")]/following-sibling::div//div[contains(@class, "text-right")]')) ) response = Selector(text=driver.page_source) data_combined.update(ScrapeData(response, x)) exporter(data_combined)
5. 验证页面结构变更
打开Chrome DevTools,检查"Realized ROI %"对应的元素结构:
- 确认列头的父元素class
- 确认数值元素的class或属性
- 确保Xpath能精准匹配所有目标数值
内容的提问来源于stack exchange,提问作者Shah Zeb
相关产品推荐
相关产品推荐

