You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Selenium获取Livescore单日全部赛事的统计数据?

解决Livescore单日赛事统计数据批量提取问题

原代码核心问题

  • 滚动过程中重复解析页面,会多次获取同一赛事URL
  • for a in full_url:是遍历URL字符串的每个字符,并非点击元素,逻辑完全错误
  • 直接在当前页面跳转赛事详情,会丢失原列表页上下文,无法继续处理剩余赛事

修正后的实现步骤

  1. 一次性收集当日所有赛事的唯一URL
    先滚动到底部加载完所有赛事,再统一解析页面收集URL,避免重复
  2. 批量处理每个赛事详情
    用新标签页打开每个赛事URL,处理完后关闭标签页,回到列表页继续
  3. 提取单场赛事统计数据
    定位赛事统计对应的HTML节点(以Livescore现有统计模块结构为例)
  4. 保存所有数据
    用字典或列表存储每场数据,最终可导出为CSV/JSON格式

修正后的代码

from selenium import webdriver
from bs4 import BeautifulSoup
import time
from urllib.parse import urljoin
import csv

# 初始化浏览器
driver = webdriver.Chrome()
driver.get('https://www.livescore.com/en/football/2022-12-01/')
time.sleep(2)

# 第一步:滚动到底部加载所有赛事
scroll_pause_time = 1
screen_height = driver.execute_script("return window.screen.height;")
i = 0
while True:
    driver.execute_script(f"window.scrollTo(0, {screen_height}*{i});")
    i += 1
    time.sleep(scroll_pause_time)
    scroll_height = driver.execute_script("return document.body.scrollHeight;")
    if screen_height * i > scroll_height:
        break

# 第二步:收集所有唯一的赛事URL
soup = BeautifulSoup(driver.page_source, 'lxml')
divs = soup.find_all('div', class_='wk Ak')
base_url = 'https://www.livescore.com'
match_urls = set()  # 用集合自动去重

for p in divs:
    url_tag = p.find('a', class_='bi')
    if url_tag and 'href' in url_tag.attrs:
        full_url = urljoin(base_url, url_tag['href'])
        match_urls.add(full_url)

# 第三步:遍历每个赛事URL,提取统计数据
all_match_data = []
main_window = driver.current_window_handle  # 保存主窗口句柄

for match_url in match_urls:
    # 新开标签页打开赛事详情
    driver.execute_script(f"window.open('{match_url}');")
    # 切换到新标签页
    driver.switch_to.window(driver.window_handles[-1])
    time.sleep(2)  # 等待页面加载完成
    
    # 提取赛事基本信息
    soup_match = BeautifulSoup(driver.page_source, 'lxml')
    match_title = soup_match.find('h1', class_='ip ip--title ip--s').get_text(strip=True) if soup_match.find('h1', class_='ip ip--title ip--s') else '未知赛事'
    
    # 提取统计数据(控球率、射门等,需根据页面实时结构调整类名)
    stats = {}
    stat_rows = soup_match.find_all('div', class_='sm__row')
    for row in stat_rows:
        stat_name = row.find('div', class_='sm__title').get_text(strip=True) if row.find('div', class_='sm__title') else ''
        home_val = row.find('div', class_='sm__value sm__value--home').get_text(strip=True) if row.find('div', class_='sm__value sm__value--home') else ''
        away_val = row.find('div', class_='sm__value sm__value--away').get_text(strip=True) if row.find('div', class_='sm__value sm__value--away') else ''
        if stat_name:
            stats[stat_name] = {'主队': home_val, '客队': away_val}
    
    # 封装单场数据
    match_data = {
        '赛事标题': match_title,
        '赛事链接': match_url,
        '统计数据': stats
    }
    all_match_data.append(match_data)
    print(f"已处理赛事:{match_title}")
    
    # 关闭当前标签页,回到主窗口
    driver.close()
    driver.switch_to.window(main_window)

# 第四步:保存数据到CSV文件
with open('livescore赛事统计.csv', 'w', newline='', encoding='utf-8') as f:
    writer = csv.writer(f)
    # 写入表头
    writer.writerow(['赛事标题', '赛事链接', '统计项', '主队数据', '客队数据'])
    # 写入每条数据
    for match in all_match_data:
        for stat_name, values in match['统计数据'].items():
            writer.writerow([match['赛事标题'], match['赛事链接'], stat_name, values['主队'], values['客队']])

# 关闭浏览器
driver.quit()

关键说明

  • URL去重:用set存储URL,避免滚动过程中重复收集同一赛事
  • 新标签页处理:保留主窗口,避免跳转后丢失原页面状态
  • 统计数据提取:Livescore的页面类名可能会更新,需根据实际页面结构调整sm__row、sm__title等类名
  • 稳定性优化:可将time.sleep()替换为WebDriverWait显式等待,等待统计模块加载完成后再提取数据

内容的提问来源于stack exchange,提问作者Anu Oyedele

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 13:00:46