如何用Beautiful Soup和Requests获取点击More[+]加载完全部足球赛事后的页面HTML
Forebet全量足球赛事数据爬取解决方案
你当前使用requests直接发起静态请求的方式只能获取页面首次加载的部分赛事数据,底部「More[+]」加载更多属于前端动态AJAX请求触发的内容,静态请求无法获取这部分数据,你可以选择以下两种方案解决:
方案1:逆向分页接口(性能最优)
- 打开浏览器开发者工具的「网络」面板,点击「More[+]」按钮即可捕获到分页请求接口
- 分析接口的请求参数(通常包含页码、每页返回条数等字段),循环调用接口直到返回数据为空,即可拿到全量赛事数据
- 该方案无需启动浏览器,运行速度最快,资源占用最低
方案2:Selenium模拟点击(改动最小,适配现有代码)
直接模拟用户操作浏览器多次点击加载按钮,直到全部内容加载完成后再提取页面数据,修改后的代码如下:
from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time leagues = {"EPL","UCL","Es1","De1","Fr1","Pt1","It1","UEL"} class ForeBet: # 获取指定联赛的所有赛事和概率,返回字符串列表 # 每条数据格式:联赛|日期|时间|主队|客队|主胜概率|平局概率|客胜概率 def get_games_and_probs(self): # 初始化浏览器驱动,这里以Chrome为例,需要提前安装对应版本的ChromeDriver driver = webdriver.Chrome() # 注意你原代码中的URL拼写错误,少了末尾的s driver.get('https://www.forebet.com/en/football-predictions') # 隐式等待设置,避免元素未加载完成报错 wait = WebDriverWait(driver, 10) # 循环点击More[+]按钮直到按钮消失 while True: try: more_btn = wait.until(EC.element_to_be_clickable((By.XPATH, "//span[contains(text(),'More') and contains(text(),'+')]"))) # 滚动到按钮位置 driver.execute_script("arguments[0].scrollIntoView();", more_btn) more_btn.click() # 等待加载完成,可根据网络情况调整等待时长 time.sleep(2) except: # 没有更多按钮时跳出循环 break # 拿到加载完成后的全量HTML内容 full_html = driver.page_source driver.quit() soup = BeautifulSoup(full_html, 'html.parser') results = list() games = soup.findAll(class_='rcnt tr_0') + soup.findAll(class_='rcnt tr_1') for game in games: short_tag = game.find(class_='shortTag').text.strip() if short_tag in leagues: date_info = game.find(class_='date_bah').text.split(" ") game_info = f"{short_tag}|{date_info[0]}|{date_info[1]}|{game.find(class_='homeTeam').text}|{game.find(class_='awayTeam').text}|{game.find(class_='fprc').findNext().text}|{game.find(class_='fprc').findNext().findNext().text}|{game.find(class_='fprc').findNext().findNext().findNext().text}" print(game_info) results.append(game_info) return results
注意事项
- 使用Selenium需要提前安装对应浏览器的驱动,Chrome驱动可匹配本地Chrome版本下载后配置到环境变量,或者在初始化
webdriver.Chrome()时指定驱动路径 - 可适当调整点击后的等待时间,避免加载未完成就进行下一次点击
- 频繁请求可能触发站点反爬策略,可适当增加每次点击的间隔时间
内容的提问来源于stack exchange,提问作者Cleto
相关产品推荐
相关产品推荐

