You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup提取match-info类a标签href无结果,求解决方案

解决Understat网页match-info链接提取失败的思路

1. 补全缺失的依赖库导入

你的代码里调用了requests.get()但未导入requests库,这会直接导致请求失败,先补上导入语句:

import requests  # 新增这行
import warnings
import numpy as np
from datetime import datetime
import json
from bs4 import BeautifulSoup

warnings.filterwarnings('ignore')

url = "https://understat.com/league/EPL/2022"
response = requests.get(url)
# ... 后续代码不变

2. 处理动态渲染的页面内容

Understat的比赛链接是通过JavaScript动态渲染到页面的,直接用BeautifulSoup抓取静态HTML无法获取到这些元素,有两种可行方案:

方案A:直接解析页面内嵌的JSON数据

页面源代码中包含window.__INITIAL_STATE__变量,里面存储了所有比赛的完整数据,直接提取并解析这个JSON就能拿到match的href:

import requests
import json
from bs4 import BeautifulSoup

url = "https://understat.com/league/EPL/2022"
response = requests.get(url)
soup = BeautifulSoup(response.content, "html.parser")

# 提取页面中的INITIAL_STATE数据
script_tag = soup.find('script', text=lambda t: '__INITIAL_STATE__' in t)
raw_data = script_tag.text.split('window.__INITIAL_STATE__ = ')[1].rstrip(';')
data = json.loads(raw_data)

# 遍历所有比赛,提取href
for match in data['dates']:
    if 'matches' in match:
        for m in match['matches']:
            print(f"match/{m['id']}")

方案B:使用浏览器渲染工具加载页面

用Selenium模拟浏览器加载页面,等待JavaScript渲染完成后再抓取DOM:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

url = "https://understat.com/league/EPL/2022"
driver = webdriver.Chrome()  # 需要提前安装对应版本的ChromeDriver
driver.get(url)

# 等待元素加载完成
WebDriverWait(driver, 10).until(
    EC.presence_of_all_elements_located((By.CLASS_NAME, "match-info"))
)

# 提取所有match-info的href
links = driver.find_elements(By.CLASS_NAME, "match-info")
for link in links:
    href = link.get_attribute("href")
    print(href.split('understat.com/')[-1])  # 只输出match/xxx部分

driver.quit()

3. 优化选择器精度

可以通过同时指定class和data-isresult属性来缩小选择范围,确保只匹配目标元素:

# 修改find_all的参数,同时过滤class和data-isresult属性
for link in soup.find_all("a", {"class": "match-info", "data-isresult": "true"}):
    href = link.get("href")
    print(href)

注:这个方法仅在静态HTML中存在该元素时有效,若元素是动态渲染的,还是需要用前面的动态渲染处理方案。

4. 添加请求头模拟浏览器访问

部分网站会拦截无标识的请求,添加User-Agent等请求头可以避免被拦截:

headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
}
response = requests.get(url, headers=headers)

内容的提问来源于stack exchange,提问作者Paul Corcoran

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 18:23:17