Python Selenium中点击加载更多后如何获取完整动态HTML?
爬取英超赛事结果仅获取200场的问题解决
问题背景
爬取https://www.skysports.com/premier-league-results/2022-23的2022-23赛季英超赛事结果,目标获取全部380场,但目前仅能抓取200场。页面存在「显示更多」按钮,点击后虽触发加载,但获取页面HTML时仍只返回点击前的内容。
尝试过使用WebDriverWait等待元素,但误以为点击后没有生成新元素,当前代码如下:
import requests import pandas as pd import time from datetime import datetime from seleniumwire import webdriver from selenium.webdriver.common.by import By from bs4 import BeautifulSoup from urllib.request import urlopen from selenium import webdriver import time import csv from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC driver = webdriver.Chrome() url = "https://www.skysports.com/premier-league-results/2022-23" driver.get(url) time.sleep(10) driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") link = driver.find_element(By.XPATH, value='//*[@id="widgetLite-9"]/button') link.click() source = driver.execute_script("return document.body.innerHTML") source = driver.page_source soup = BeautifulSoup(source,'lxml') div_tags = soup.find_all('div', attrs={'class': 'fixres__body'}) soup = div_tags[0] matches = [] date_tags = soup.find_all('h4', class_='fixres__header2') for tag in tags: if tag.name == 'h4': date = tag.text.strip() elif tag.name == 'div' and tag['class'] == ['fixres__item']: team1 = tag.find('span', class_='matches__participant--side1').text.strip() team2 = tag.find('span', class_='matches__participant--side2').text.strip() match_time = tag.find('span', class_='matches__date').text.strip() matches.append([date, match_time, team1, team2]) with open('matches.csv', 'w', newline='') as csvfile: writer = csv.writer(csvfile) writer.writerow(['Date', 'Time', 'Team 1', 'Team 2']) # Write the header row writer.writerows(matches) print(len(matches))
问题原因
- 未等待异步加载完成:点击「显示更多」后,页面通过AJAX异步加载新数据,直接获取
page_source会拿到未更新的DOM内容。 - 「显示更多」需多次点击:该按钮不会一次性加载所有剩余赛事,需重复点击直到按钮消失,才能获取全部380场数据。
- 遍历逻辑错误:代码中定义了
date_tags却遍历未定义的tags,且未同时处理日期标签和赛事条目标签,导致部分数据未被抓取。
修复方案
import time import csv from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup driver = webdriver.Chrome() url = "https://www.skysports.com/premier-league-results/2022-23" driver.get(url) # 循环点击「显示更多」直到按钮消失 while True: try: # 滚动到按钮位置并等待可点击 show_more_btn = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.XPATH, '//*[@id="widgetLite-9"]/button')) ) driver.execute_script("arguments[0].scrollIntoView();", show_more_btn) show_more_btn.click() # 等待新内容加载(等待赛事条目数量增加) WebDriverWait(driver, 10).until( lambda d: len(d.find_elements(By.CLASS_NAME, 'fixres__item')) > 200 ) time.sleep(2) # 额外缓冲时间 except: # 按钮不存在时退出循环 break # 获取完整页面源码 source = driver.page_source soup = BeautifulSoup(source, 'lxml') fixres_body = soup.find('div', class_='fixres__body') matches = [] current_date = "" # 遍历所有日期和赛事标签 for tag in fixres_body.find_all(['h4', 'div']): if tag.name == 'h4' and 'fixres__header2' in tag.get('class', []): current_date = tag.text.strip() elif tag.name == 'div' and 'fixres__item' in tag.get('class', []): team1 = tag.find('span', class_='matches__participant--side1').text.strip() team2 = tag.find('span', class_='matches__participant--side2').text.strip() match_time = tag.find('span', class_='matches__date').text.strip() matches.append([current_date, match_time, team1, team2]) # 写入CSV文件 with open('matches.csv', 'w', newline='', encoding='utf-8') as csvfile: writer = csv.writer(csvfile) writer.writerow(['Date', 'Time', 'Team 1', 'Team 2']) writer.writerows(matches) print(f"共抓取{len(matches)}场赛事") driver.quit()
关键改进点
- 循环点击加载:通过
while循环重复点击「显示更多」按钮,直到按钮无法找到(所有内容加载完成)。 - 显式等待加载:使用
WebDriverWait等待按钮可点击、等待赛事条目数量增加,确保新内容加载完成后再继续操作。 - 修正遍历逻辑:同时遍历日期标签和赛事条目标签,正确关联每场赛事对应的日期。
- 编码处理:写入CSV时指定
encoding='utf-8',避免中文乱码。
内容的提问来源于stack exchange,提问作者Vinay
相关产品推荐
相关产品推荐

