You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Selenium中点击加载更多后如何获取完整动态HTML?

爬取英超赛事结果仅获取200场的问题解决

问题背景

爬取https://www.skysports.com/premier-league-results/2022-23的2022-23赛季英超赛事结果,目标获取全部380场,但目前仅能抓取200场。页面存在「显示更多」按钮,点击后虽触发加载,但获取页面HTML时仍只返回点击前的内容。

尝试过使用WebDriverWait等待元素,但误以为点击后没有生成新元素,当前代码如下:

import requests
import pandas as pd
import time
from datetime import datetime
from seleniumwire import webdriver 
from selenium.webdriver.common.by import By
from bs4 import BeautifulSoup
from urllib.request import urlopen
from selenium import webdriver
import time
import csv
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

driver = webdriver.Chrome()        
url = "https://www.skysports.com/premier-league-results/2022-23"
driver.get(url)
time.sleep(10)
driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")

link = driver.find_element(By.XPATH, value='//*[@id="widgetLite-9"]/button')
link.click()

source = driver.execute_script("return document.body.innerHTML")    
source = driver.page_source
soup = BeautifulSoup(source,'lxml')
div_tags = soup.find_all('div', attrs={'class': 'fixres__body'})
soup = div_tags[0]

matches = []
date_tags = soup.find_all('h4', class_='fixres__header2')
for tag in tags:
    if tag.name == 'h4':
        date = tag.text.strip()
    elif tag.name == 'div' and tag['class'] == ['fixres__item']:
        team1 = tag.find('span', class_='matches__participant--side1').text.strip()
        team2 = tag.find('span', class_='matches__participant--side2').text.strip()
        match_time = tag.find('span', class_='matches__date').text.strip()      
        matches.append([date, match_time, team1, team2])
        
with open('matches.csv', 'w', newline='') as csvfile:
    writer = csv.writer(csvfile)
    writer.writerow(['Date', 'Time', 'Team 1', 'Team 2'])  # Write the header row
    writer.writerows(matches)

print(len(matches))

问题原因

  1. 未等待异步加载完成:点击「显示更多」后,页面通过AJAX异步加载新数据,直接获取page_source会拿到未更新的DOM内容。
  2. 「显示更多」需多次点击:该按钮不会一次性加载所有剩余赛事,需重复点击直到按钮消失,才能获取全部380场数据。
  3. 遍历逻辑错误:代码中定义了date_tags却遍历未定义的tags,且未同时处理日期标签和赛事条目标签,导致部分数据未被抓取。

修复方案

import time
import csv
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup

driver = webdriver.Chrome()        
url = "https://www.skysports.com/premier-league-results/2022-23"
driver.get(url)

# 循环点击「显示更多」直到按钮消失
while True:
    try:
        # 滚动到按钮位置并等待可点击
        show_more_btn = WebDriverWait(driver, 10).until(
            EC.element_to_be_clickable((By.XPATH, '//*[@id="widgetLite-9"]/button'))
        )
        driver.execute_script("arguments[0].scrollIntoView();", show_more_btn)
        show_more_btn.click()
        # 等待新内容加载(等待赛事条目数量增加)
        WebDriverWait(driver, 10).until(
            lambda d: len(d.find_elements(By.CLASS_NAME, 'fixres__item')) > 200
        )
        time.sleep(2)  # 额外缓冲时间
    except:
        # 按钮不存在时退出循环
        break

# 获取完整页面源码
source = driver.page_source
soup = BeautifulSoup(source, 'lxml')
fixres_body = soup.find('div', class_='fixres__body')

matches = []
current_date = ""
# 遍历所有日期和赛事标签
for tag in fixres_body.find_all(['h4', 'div']):
    if tag.name == 'h4' and 'fixres__header2' in tag.get('class', []):
        current_date = tag.text.strip()
    elif tag.name == 'div' and 'fixres__item' in tag.get('class', []):
        team1 = tag.find('span', class_='matches__participant--side1').text.strip()
        team2 = tag.find('span', class_='matches__participant--side2').text.strip()
        match_time = tag.find('span', class_='matches__date').text.strip()      
        matches.append([current_date, match_time, team1, team2])

# 写入CSV文件
with open('matches.csv', 'w', newline='', encoding='utf-8') as csvfile:
    writer = csv.writer(csvfile)
    writer.writerow(['Date', 'Time', 'Team 1', 'Team 2'])
    writer.writerows(matches)

print(f"共抓取{len(matches)}场赛事")
driver.quit()

关键改进点

  • 循环点击加载:通过while循环重复点击「显示更多」按钮,直到按钮无法找到(所有内容加载完成)。
  • 显式等待加载:使用WebDriverWait等待按钮可点击、等待赛事条目数量增加,确保新内容加载完成后再继续操作。
  • 修正遍历逻辑:同时遍历日期标签和赛事条目标签,正确关联每场赛事对应的日期。
  • 编码处理:写入CSV时指定encoding='utf-8',避免中文乱码。

内容的提问来源于stack exchange,提问作者Vinay

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 13:42:51