You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium与ChromeDriver爬虫语法错误排查及代码正确性验证

英超官网Selenium爬虫问题排查与修正

一、语法错误解决

你遇到的SyntaxError是因为XPATH字符串里用了HTML转义字符",Python无法识别,替换为Python支持的引号格式即可:

# 错误写法
matches = driver.find_element(By.XPATH,'//*[@id="mainContent"]/div[3]/div[1]/div[2]/section/div[1]/ul/li[1]')

# 正确写法(单引号包裹XPATH,内部双引号直接使用)
matches = driver.find_element(By.XPATH,'//*[@id="mainContent"]/div[3]/div[1]/div[2]/section/div[1]/ul/li[1]')

二、代码逻辑问题修正(无法爬取所有赛季数据)

你的代码存在多处逻辑错误,既无法正确获取单赛季数据,更无法覆盖所有赛季,以下是问题点和修正方案:

1. WebDriverWait用法错误

原代码中WebDriverWait的语法不符合规范,且元素定位逻辑混乱(用By.ID却传入XPATH表达式),修正为:

from selenium.webdriver.support import expected_conditions as EC

# 等待关闭广告按钮可点击后执行点击
WebDriverWait(driver, 20).until(EC.element_to_be_clickable((By.ID, "advertClose"))).click()

2. 变量覆盖与循环逻辑错误

  • 原代码先将matches赋值为单个元素,随后又赋值为空列表,导致后续循环完全无法执行
  • 循环中错误使用matches.find_element(应使用当前循环项match),且XPATH用绝对路径会重复获取第一个元素,需改为相对路径(以./开头)

3. 无法爬取所有赛季的核心问题

当前代码仅打开了单个赛季的链接(se=363对应特定赛季),要爬取所有赛季,需添加赛季切换逻辑:找到赛季下拉框,遍历所有选项并依次切换爬取。

修正后的完整示例代码

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.service import Service
from webdriver_manager.chrome import ChromeDriverManager
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from time import sleep

driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()))
# 打开英超战绩主页面
driver.get("https://www.premierleague.com/results")
sleep(3)

# 关闭初始弹窗(容错处理,避免弹窗不存在报错)
try:
    WebDriverWait(driver, 10).until(EC.element_to_be_clickable((By.XPATH, '//*[@class="_24Il51SkQ29P1pCkJOUO-7"]/button'))).click()
except:
    pass

# 关闭广告弹窗(容错处理)
try:
    WebDriverWait(driver, 10).until(EC.element_to_be_clickable((By.ID, "advertClose"))).click()
except:
    pass

# 点击展开赛季下拉框
season_dropdown = WebDriverWait(driver, 10).until(EC.element_to_be_clickable((By.XPATH, '//*[@id="mainContent"]/div[3]/div[1]/div[1]/div/div[2]')))
season_dropdown.click()
sleep(2)
# 获取所有赛季选项
season_options = WebDriverWait(driver, 10).until(EC.presence_of_all_elements_located((By.XPATH, '//*[@id="mainContent"]/div[3]/div[1]/div[1]/div/div[2]/ul/li')))

# 遍历每个赛季爬取数据
for season in season_options:
    season_name = season.text
    print(f"\n===== 开始爬取赛季:{season_name} =====")
    season.click()
    sleep(5)
    
    # 获取当前赛季所有比赛条目
    matches = WebDriverWait(driver, 10).until(EC.presence_of_all_elements_located((By.XPATH, '//*[@id="mainContent"]/div[3]/div[1]/div[2]/section/div[1]/ul/li')))
    for match in matches:
        try:
            # 用相对路径获取两队名称和比分
            team_elements = match.find_elements(By.XPATH, './div/span')
            team1 = team_elements[0].text
            team2 = team_elements[2].text
            score = match.find_element(By.XPATH, './div/span/span[1]/span[2]').text
            print(f"{team1} {score} {team2}")
        except Exception as e:
            print(f"单场数据获取失败:{str(e)}")
    
    # 重新展开赛季下拉框,准备切换下一个赛季
    season_dropdown = WebDriverWait(driver, 10).until(EC.element_to_be_clickable((By.XPATH, '//*[@id="mainContent"]/div[3]/div[1]/div[1]/div/div[2]')))
    season_dropdown.click()
    sleep(2)

driver.quit()

注意事项

  • 英超官网有反爬机制,需合理设置等待时间,避免频繁操作触发限制
  • 页面元素可能随网站更新变化,需定期检查XPATH定位是否有效

内容的提问来源于stack exchange,提问作者Jarod

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 20:10:26