如何从Flashscore页面爬取所有比赛详情的超链接?
解决Flashscore爬取比赛详情链接的问题
Flashscore的比赛数据是动态加载的——你用requests.get()获取的只是页面的初始静态HTML,而比赛列表和详情链接是页面加载完成后通过JavaScript渲染出来的,所以BeautifulSoup解析静态内容时根本找不到这些目标链接。
下面提供两种可行的解决方法:
方案1:用Selenium模拟浏览器加载页面
Selenium可以模拟真实浏览器的行为,等待页面完全渲染(包括JS加载的内容),之后再解析页面元素。
from selenium import webdriver from selenium.webdriver.chrome.options import Options from bs4 import BeautifulSoup import time # 设置Chrome无头模式(可选,不弹出浏览器窗口) chrome_options = Options() chrome_options.add_argument("--headless=new") # 初始化浏览器驱动(需提前下载对应版本的chromedriver) driver = webdriver.Chrome(options=chrome_options) url = 'https://www.flashscore.com/football/france/ligue-1-2022-2023/results/' driver.get(url) # 等待页面加载完成(可根据网络情况调整等待时间) time.sleep(3) # 获取渲染后的页面源码 page_source = driver.page_source soup = BeautifulSoup(page_source, 'html.parser') # 提取比赛详情链接 match_links = [] # 比赛链接的a标签通常带有特定class:event__link for link in soup.find_all('a', class_='event__link'): href = link.get('href') if href and '/match/' in href: full_link = f"https://www.flashscore.com{href}" match_links.append(full_link) print(full_link) # 关闭浏览器 driver.quit()
注意:需要提前下载对应浏览器版本的驱动(比如Chrome的chromedriver),并确保驱动路径配置正确。
方案2:直接调用网站的API接口
Flashscore会通过XHR请求获取比赛数据,你可以通过浏览器开发者工具找到对应的API接口,直接请求接口获取JSON数据,再从中提取比赛ID构造详情链接。
操作步骤:
- 打开Flashscore结果页面,按F12打开开发者工具,切换到Network标签;
- 刷新页面,筛选XHR类型的请求,找到返回比赛数据的接口(通常地址类似
https://d.flashscore.com/x/feed/d_hh_[联赛标识]_en_1); - 查看响应内容,提取其中的比赛
id字段; - 用
requests请求接口并构造详情链接:
import requests import json # 替换为你抓取到的API接口地址 api_url = 'https://d.flashscore.com/x/feed/d_hh_igz8y_en_1' headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Referer': 'https://www.flashscore.com/football/france/ligue-1-2022-2023/results/' } response = requests.get(api_url, headers=headers) # 解析Flashscore特殊格式的API响应 data = json.loads(response.text.split('(')[1].split(')')[0]) match_links = [] # 遍历比赛数据提取ID并构造链接 for event in data['doc'][0]['data']['events']: match_id = event['id'] full_link = f"https://www.flashscore.com/match/{match_id}/#/match-summary/match-summary" match_links.append(full_link) print(full_link)
注意:API接口地址可能随网站更新变化,需要重新抓取;请求时需带上合适的headers,避免被网站拦截。
内容的提问来源于stack exchange,提问作者mzietal
相关产品推荐
相关产品推荐

