使用Python BeautifulSoup抓取FotMob赛事链接返回空列表求助
解决FotMob赛程链接抓取返回空列表的问题
问题原因
- 你用
requests获取的是页面初始静态HTML,但目标链接是通过JavaScript动态渲染生成的,静态源码中不存在这些标签,因此BeautifulSoup无法找到对应元素。 - 你依赖的
css-11smhdq-FtContainer e1ym2d3s2是动态生成的类名,这类名称会随网站更新频繁变化,可靠性极低。
解决方案
方案一:用Selenium模拟浏览器加载动态内容
这种方法会真实模拟浏览器加载页面,等待动态内容渲染完成后再抓取:
- 先安装依赖:
pip install selenium
- 下载对应浏览器的驱动(比如ChromeDriver,需和浏览器版本匹配),然后编写代码:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup url = "https://www.fotmob.com/teams/8676/fixtures/wycombe-wanderers?page=1" # 初始化Chrome驱动(确保驱动路径正确,或已配置到环境变量) driver = webdriver.Chrome() driver.get(url) # 等待目标元素加载完成(最多等待10秒) wait = WebDriverWait(driver, 10) wait.until(EC.presence_of_element_located((By.CLASS_NAME, "e1ym2d3s2"))) # 获取渲染后的完整页面源码 html_doc = driver.page_source soup = BeautifulSoup(html_doc, 'html.parser') # 用类名中的固定部分匹配(避免动态类名变化) links = [tag['href'] for tag in soup.find_all('a', class_=lambda cls: cls and 'FtContainer' in cls)] print(links) # 关闭浏览器 driver.quit()
方案二:直接调用网站API接口(更高效稳定)
动态网站的内容通常来自API接口,通过浏览器开发者工具的「网络」面板可以找到对应的接口,直接请求JSON数据更高效:
import requests # 抓取到的FotMob球队赛程API接口 api_url = "https://www.fotmob.com/api/teams?id=8676&season=2024&fixtures=true" # 添加请求头模拟浏览器,避免被拦截 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" } response = requests.get(api_url, headers=headers) data = response.json() # 从JSON数据中提取赛程ID,构造完整链接 fixtures = data['fixtures']['allFixtures'] match_links = [f"https://www.fotmob.com/match/{item['id']}" for item in fixtures] print(match_links)
内容的提问来源于stack exchange,提问作者Ben303
相关产品推荐
相关产品推荐

