You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup提取英超赛事页面href标签无返回结果求助

问题原因与解决方法

英超官网的赛事结果页面采用动态渲染,直接用BeautifulSoup请求静态HTML无法获取到赛事链接——这些内容是页面加载后通过JavaScript异步加载的,静态源码里根本没有对应的标签。以下是两种可行的解决方案:

方案1:结合Selenium获取动态渲染页面

用Selenium模拟浏览器加载完整页面,再用BeautifulSoup解析渲染后的源码:

from selenium import webdriver
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
from bs4 import BeautifulSoup

# 初始化Chrome浏览器(需提前安装对应版本的chromedriver)
driver = webdriver.Chrome()
driver.get("https://www.premierleague.com/results")

# 等待赛事列表加载完成(最多等10秒)
try:
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CSS_SELECTOR, "a.matchFixtureContainer"))
    )
finally:
    page_source = driver.page_source
    driver.quit()

# 解析页面提取链接
soup = BeautifulSoup(page_source, "html.parser")
match_links = []
for a_tag in soup.select("a.matchFixtureContainer"):
    href = a_tag.get("href")
    if href and "/match/" in href:
        # 统一格式为示例中的相对链接或拼接完整链接
        match_links.append(href if href.startswith("//") else f"//www.premierleague.com{href}")

print(match_links)

方案2:直接解析页面内嵌的JSON数据

很多动态网站会把核心数据存在页面的script标签里,无需模拟浏览器即可提取:

import requests
import json
from bs4 import BeautifulSoup

headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}
response = requests.get("https://www.premierleague.com/results", headers=headers)
soup = BeautifulSoup(response.text, "html.parser")

# 定位存储赛事数据的script标签
for script in soup.find_all("script"):
    if script.string and "window.__INITIAL_STATE__" in script.string:
        # 提取并解析JSON数据
        json_raw = script.string.split("window.__INITIAL_STATE__ = ")[1].split(";")[0].strip()
        data = json.loads(json_raw)
        # 从JSON中提取赛事ID并拼接链接
        match_list = data.get("matches", {}).get("matches", [])
        match_links = [f"//www.premierleague.com/match/{match['id']}" for match in match_list if "id" in match]
        print(match_links)
        break

注意事项

  • 英超官网有反爬机制,请求时务必带上合法的User-Agent,避免频繁请求导致IP被封禁。
  • 方案2中的JSON结构可能随网站更新变动,若失效可打开浏览器开发者工具(F12)查看最新的数据存储位置。

内容的提问来源于stack exchange,提问作者Paul Corcoran

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 10:25:14