从MLS赛事页面提取比赛链接href的技术实现问题
解决MLS赛事页面提取比赛链接的问题
方案一:Selenium渲染后用BeautifulSoup提取
你的思路是对的,动态加载的内容需要等页面渲染完成后再提取。以下是补充完整的代码,包含链接提取和翻页逻辑:
import requests from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import TimeoutException, NoSuchElementException from selenium.webdriver.common.by import By from time import sleep import pandas as pd import warnings import numpy as np from datetime import datetime import json from bs4 import BeautifulSoup warnings.filterwarnings('ignore') base_url = 'https://www.mlssoccer.com/schedule/scores#competition=mls-regular-season&club=all&date=2023-02-20' urls = [] option = Options() option.headless = False driver = webdriver.Chrome("##########", options=option) driver.get(base_url) # 处理cookie弹窗(改用相对定位,避免绝对XPath失效) try: WebDriverWait(driver, 15).until( EC.element_to_be_clickable((By.CSS_SELECTOR, 'button[aria-label="Accept All Cookies"]')) ).click() except TimeoutException: print("未检测到cookie弹窗,继续执行") # 循环处理翻页 while True: # 等待比赛列表区域加载完成 try: WebDriverWait(driver, 20).until( EC.presence_of_element_located((By.CLASS_NAME, "mls-l-module--match-list")) ) except TimeoutException: print("比赛列表加载超时,跳过当前页") break # 获取渲染后的页面源码,交给BeautifulSoup解析 page_source = driver.page_source soup = BeautifulSoup(page_source, 'html.parser') # 定位目标section并提取链接 match_section = soup.find('section', class_='mls-l-module mls-l-module--match-list') if match_section: # 提取所有包含比赛详情的a标签(MLS比赛链接通常以/matches/开头) match_links = match_section.find_all('a', href=lambda x: x and '/matches/' in x) for link in match_links: full_url = f"https://www.mlssoccer.com{link['href']}" if full_url not in urls: urls.append(full_url) # 尝试点击下一页 try: next_btn = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.XPATH, '//button[contains(@class, "pagination__next")]')) ) # 判断按钮是否可用(有些网站会给禁用按钮加disabled属性) if 'disabled' in next_btn.get_attribute('class'): print("已到最后一页,停止翻页") break next_btn.click() # 等待页面更新,也可以用显式等待新的比赛元素出现 sleep(2) except TimeoutException: print("未找到下一页按钮,停止翻页") break # 输出结果 print(f"共提取到{len(urls)}个比赛链接:") for url in urls: print(url) driver.quit()
关键说明:
- 改用相对定位处理元素(比如cookie按钮、下一页按钮),避免绝对XPath因页面结构变化失效
- 用
lambda x: x and '/matches/' in x过滤比赛链接,确保只提取目标href - 翻页时判断按钮是否禁用,避免无限循环
方案二:直接调用API(更高效)
动态加载的内容通常来自后端API,直接请求API比模拟浏览器更快速、资源占用更少。步骤如下:
- 打开浏览器开发者工具(F12),切换到Network标签
- 刷新MLS赛事页面,搜索包含
schedule或match的XHR请求 - 找到返回比赛数据的API接口(通常是GET请求,返回JSON格式)
以下是示例代码(需替换为实际找到的API地址):
import requests headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } # 替换为实际找到的API URL,参数可根据需要调整(比如日期、赛事类型) api_url = "https://api.mlssoccer.com/v1/schedule?competition=mls-regular-season&club=all&date=2023-02-20" response = requests.get(api_url, headers=headers) response.raise_for_status() # 检查请求是否成功 match_data = response.json() urls = [] # 遍历返回的比赛数据,拼接链接(具体字段需看API返回的JSON结构) for match in match_data.get('matches', []): match_slug = match.get('slug') if match_slug: full_url = f"https://www.mlssoccer.com/matches/{match_slug}" urls.append(full_url) print(f"共提取到{len(urls)}个比赛链接:") for url in urls: print(url)
关键说明:
- 需根据实际API返回的JSON结构调整字段(比如有的API用
id而非slug) - 添加
User-Agent头模拟浏览器,避免被API拦截
内容的提问来源于stack exchange,提问作者Paul Corcoran
相关产品推荐
相关产品推荐

