You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于两个用户输入实现网页爬取并存储指定字段为独立变量

英超赛程爬取代码优化方案

核心修改点

  • 新增赛事匹配逻辑:遍历赛程表格行,对比用户输入的比赛轮次与主队
  • 实现目标字段提取:从匹配到的行中提取日期、开赛时间、客队、场馆并存储为独立变量
  • 增加输入容错:将主队名称转为小写对比,避免大小写差异导致的匹配失败

优化后完整代码

import requests
from bs4 import BeautifulSoup

# 获取用户输入
gw = input("Enter the gameweek of the game you would like to attend \n")
hometeam = input("Enter the name of the home team \n")

# 请求并解析赛程页面
url = requests.get("https://fbref.com/en/comps/9/schedule/Premier-League-Scores-and-Fixtures") 
fixtures_soup = BeautifulSoup(url.content, "html.parser") 
schedule_table = fixtures_soup.find('table', id="sched_2023-2024_9_1")

# 初始化目标变量
match_date = None
kickoff_time = None
away_team = None
stadium = None

# 遍历所有赛事行(跳过表头)
for row in schedule_table.find_all('tr')[1:]:
    current_gw = row.find('th', {'data-stat': 'gameweek'}).text.strip()
    current_home = row.find('td', {'data-stat': 'home_team'}).text.strip()
    
    # 匹配用户指定的轮次和主队
    if current_gw == gw and current_home.lower() == hometeam.lower():
        # 提取对应字段
        match_date = row.find('td', {'data-stat': 'date'}).text.strip()
        kickoff_time = row.find('td', {'data-stat': 'time'}).text.strip()
        away_team = row.find('td', {'data-stat': 'away_team'}).text.strip()
        stadium = row.find('td', {'data-stat': 'venue'}).text.strip()
        break

# 输出结果
if match_date:
    print(f"赛事日期: {match_date}")
    print(f"开赛时间: {kickoff_time}")
    print(f"客队: {away_team}")
    print(f"场馆: {stadium}")
else:
    print("未找到符合条件的赛事")

代码说明

  • 页面解析:通过表格ID定位赛程表,确保只处理目标数据区域
  • 匹配逻辑:逐行检查赛事的轮次和主队,匹配成功后立即停止遍历提升效率
  • 字段提取:利用HTML标签的data-stat属性精准定位目标单元格,避免结构变动导致的失效

内容的提问来源于stack exchange,提问作者hdhdicb

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 12:03:24