如何用Selenium获取Livescore单日全部赛事的统计数据?
解决Livescore单日赛事统计数据批量提取问题
原代码核心问题
- 滚动过程中重复解析页面,会多次获取同一赛事URL
for a in full_url:是遍历URL字符串的每个字符,并非点击元素,逻辑完全错误- 直接在当前页面跳转赛事详情,会丢失原列表页上下文,无法继续处理剩余赛事
修正后的实现步骤
- 一次性收集当日所有赛事的唯一URL
先滚动到底部加载完所有赛事,再统一解析页面收集URL,避免重复 - 批量处理每个赛事详情
用新标签页打开每个赛事URL,处理完后关闭标签页,回到列表页继续 - 提取单场赛事统计数据
定位赛事统计对应的HTML节点(以Livescore现有统计模块结构为例) - 保存所有数据
用字典或列表存储每场数据,最终可导出为CSV/JSON格式
修正后的代码
from selenium import webdriver from bs4 import BeautifulSoup import time from urllib.parse import urljoin import csv # 初始化浏览器 driver = webdriver.Chrome() driver.get('https://www.livescore.com/en/football/2022-12-01/') time.sleep(2) # 第一步:滚动到底部加载所有赛事 scroll_pause_time = 1 screen_height = driver.execute_script("return window.screen.height;") i = 0 while True: driver.execute_script(f"window.scrollTo(0, {screen_height}*{i});") i += 1 time.sleep(scroll_pause_time) scroll_height = driver.execute_script("return document.body.scrollHeight;") if screen_height * i > scroll_height: break # 第二步:收集所有唯一的赛事URL soup = BeautifulSoup(driver.page_source, 'lxml') divs = soup.find_all('div', class_='wk Ak') base_url = 'https://www.livescore.com' match_urls = set() # 用集合自动去重 for p in divs: url_tag = p.find('a', class_='bi') if url_tag and 'href' in url_tag.attrs: full_url = urljoin(base_url, url_tag['href']) match_urls.add(full_url) # 第三步:遍历每个赛事URL,提取统计数据 all_match_data = [] main_window = driver.current_window_handle # 保存主窗口句柄 for match_url in match_urls: # 新开标签页打开赛事详情 driver.execute_script(f"window.open('{match_url}');") # 切换到新标签页 driver.switch_to.window(driver.window_handles[-1]) time.sleep(2) # 等待页面加载完成 # 提取赛事基本信息 soup_match = BeautifulSoup(driver.page_source, 'lxml') match_title = soup_match.find('h1', class_='ip ip--title ip--s').get_text(strip=True) if soup_match.find('h1', class_='ip ip--title ip--s') else '未知赛事' # 提取统计数据(控球率、射门等,需根据页面实时结构调整类名) stats = {} stat_rows = soup_match.find_all('div', class_='sm__row') for row in stat_rows: stat_name = row.find('div', class_='sm__title').get_text(strip=True) if row.find('div', class_='sm__title') else '' home_val = row.find('div', class_='sm__value sm__value--home').get_text(strip=True) if row.find('div', class_='sm__value sm__value--home') else '' away_val = row.find('div', class_='sm__value sm__value--away').get_text(strip=True) if row.find('div', class_='sm__value sm__value--away') else '' if stat_name: stats[stat_name] = {'主队': home_val, '客队': away_val} # 封装单场数据 match_data = { '赛事标题': match_title, '赛事链接': match_url, '统计数据': stats } all_match_data.append(match_data) print(f"已处理赛事:{match_title}") # 关闭当前标签页,回到主窗口 driver.close() driver.switch_to.window(main_window) # 第四步:保存数据到CSV文件 with open('livescore赛事统计.csv', 'w', newline='', encoding='utf-8') as f: writer = csv.writer(f) # 写入表头 writer.writerow(['赛事标题', '赛事链接', '统计项', '主队数据', '客队数据']) # 写入每条数据 for match in all_match_data: for stat_name, values in match['统计数据'].items(): writer.writerow([match['赛事标题'], match['赛事链接'], stat_name, values['主队'], values['客队']]) # 关闭浏览器 driver.quit()
关键说明
- URL去重:用
set存储URL,避免滚动过程中重复收集同一赛事 - 新标签页处理:保留主窗口,避免跳转后丢失原页面状态
- 统计数据提取:Livescore的页面类名可能会更新,需根据实际页面结构调整
sm__row、sm__title等类名 - 稳定性优化:可将
time.sleep()替换为WebDriverWait显式等待,等待统计模块加载完成后再提取数据
内容的提问来源于stack exchange,提问作者Anu Oyedele
相关产品推荐
相关产品推荐

