You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

XPath爬取异常求助:重复获取相同赔率值问题

解决Bet365爬虫XPath重复获取赔率的问题

你的问题核心出在XPath的定位逻辑上——用following::div会从当前参赛元素开始,往后匹配所有符合条件的节点,而第一个匹配到的总是页面里的第一个赔率(也就是你看到的重复的3.00),而不是当前参赛项对应行的赔率。

问题根源分析

原xp_bp1的写法:

.//following::div[contains(@class,'sl-MarketCouponValuesExplicit33')][./div[contains(@class,'gl-MarketColumnHeader')][.='1']]//span[@class='gl-ParticipantOddsOnly_Odds']

这个XPath会从当前elem(参赛名称元素)往后遍历所有带有sl-MarketCouponValuesExplicit33类的div,并且筛选出包含标题为"1"的列的容器,然后取第一个匹配的赔率。但页面里所有参赛项共享同一个赔率列容器,所以每次都会取到同一个初始值。

修正方案

我们需要调整定位逻辑:先找到当前参赛元素所在的行容器,再在这个容器内精准定位对应"1"列的赔率。

修正后的完整脚本

import csv
import time
from selenium import webdriver
from selenium.common.exceptions import TimeoutException, NoSuchElementException
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By

driver = webdriver.Chrome()
driver.set_window_size(1024, 600)
driver.maximize_window()
# 移除重复的页面请求
driver.get('https://www.bet365.com.au/#/AC/B1/C1/D13/E108/F16/S1/')

# 用显式等待替代固定sleep,更高效
try:
    WebDriverWait(driver, 15).until(
        EC.presence_of_element_located((By.XPATH, ".//div[contains(@class, 'sl-CouponParticipantWithBookCloses')]"))
    )
except TimeoutException:
    print("页面加载超时,退出程序")
    driver.quit()
    exit()

# 定位整个参赛项的行容器,而非单独的名称元素
groups = ".//div[contains(@class, 'sl-CouponParticipantWithBookCloses')]"
# 修正后的XPath:在当前行容器内,找到标题为"1"的列对应的赔率
xp_bp1 = ".//div[contains(@class,'sl-MarketCouponValuesExplicit33')]//div[contains(@class,'gl-MarketColumnHeader')][.='1']/following-sibling::div//span[@class='gl-ParticipantOddsOnly_Odds']"

while True:
    try:
        time.sleep(2)
        data = []
        # 等待参赛项元素加载完成
        WebDriverWait(driver, 10).until(EC.presence_of_all_elements_located((By.XPATH, groups)))
        for elem in driver.find_elements_by_xpath(groups):
            try:
                # 可选:获取参赛名称,方便对应赔率
                participant_name = elem.find_element_by_xpath(".//*[contains(@class, 'sl-CouponParticipantWithBookCloses_Name')]").text
                bp1 = elem.find_element_by_xpath(xp_bp1).text
            except NoSuchElementException:
                bp1 = None
                participant_name = None
            current_url = driver.current_url
            data.append([participant_name, bp1])
        print("获取到的赔率数据:", data)
        # 写入CSV
        with open('test.csv', 'a', newline='', encoding="utf-8") as outfile:
            writer = csv.writer(outfile)
            for row in data:
                writer.writerow(row + [current_url])
    except TimeoutException:
        print("等待元素超时,继续尝试")
        pass
    except Exception as ex:
        print(f"程序发生错误:{str(ex)}")
        break
# 退出浏览器
driver.quit()

关键修正点

  1. 调整groups定位范围:不再只定位参赛名称文本,而是定位整个参赛项的行容器sl-CouponParticipantWithBookCloses,确保我们能在当前行内查找对应赔率。
  2. 重写xp_bp1逻辑:
    • 移除全局的following::查找,改为在当前行容器内搜索
    • 用following-sibling::div精准定位标题"1"对应的下一个兄弟元素(即赔率所在的div),确保取到的是当前参赛项对应列的赔率
  3. 优化等待逻辑:用WebDriverWait显式等待替代固定time.sleep(),既提升效率又避免因页面加载慢导致的元素未找到问题。

额外提示

  • Bet365有严格的反爬机制,频繁请求或操作可能会触发限制,建议适当延长请求间隔,避免被封IP。
  • 若后续页面结构发生变化,可通过浏览器开发者工具(F12)重新定位元素的class或层级关系,调整XPath。

内容的提问来源于stack exchange,提问作者user9155788

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:03:41