如何基于两个用户输入实现网页爬取并存储指定字段为独立变量
英超赛程爬取代码优化方案
核心修改点
- 新增赛事匹配逻辑:遍历赛程表格行,对比用户输入的比赛轮次与主队
- 实现目标字段提取:从匹配到的行中提取日期、开赛时间、客队、场馆并存储为独立变量
- 增加输入容错:将主队名称转为小写对比,避免大小写差异导致的匹配失败
优化后完整代码
import requests from bs4 import BeautifulSoup # 获取用户输入 gw = input("Enter the gameweek of the game you would like to attend \n") hometeam = input("Enter the name of the home team \n") # 请求并解析赛程页面 url = requests.get("https://fbref.com/en/comps/9/schedule/Premier-League-Scores-and-Fixtures") fixtures_soup = BeautifulSoup(url.content, "html.parser") schedule_table = fixtures_soup.find('table', id="sched_2023-2024_9_1") # 初始化目标变量 match_date = None kickoff_time = None away_team = None stadium = None # 遍历所有赛事行(跳过表头) for row in schedule_table.find_all('tr')[1:]: current_gw = row.find('th', {'data-stat': 'gameweek'}).text.strip() current_home = row.find('td', {'data-stat': 'home_team'}).text.strip() # 匹配用户指定的轮次和主队 if current_gw == gw and current_home.lower() == hometeam.lower(): # 提取对应字段 match_date = row.find('td', {'data-stat': 'date'}).text.strip() kickoff_time = row.find('td', {'data-stat': 'time'}).text.strip() away_team = row.find('td', {'data-stat': 'away_team'}).text.strip() stadium = row.find('td', {'data-stat': 'venue'}).text.strip() break # 输出结果 if match_date: print(f"赛事日期: {match_date}") print(f"开赛时间: {kickoff_time}") print(f"客队: {away_team}") print(f"场馆: {stadium}") else: print("未找到符合条件的赛事")
代码说明
- 页面解析:通过表格ID定位赛程表,确保只处理目标数据区域
- 匹配逻辑:逐行检查赛事的轮次和主队,匹配成功后立即停止遍历提升效率
- 字段提取:利用HTML标签的
data-stat属性精准定位目标单元格,避免结构变动导致的失效
内容的提问来源于stack exchange,提问作者hdhdicb
相关产品推荐
相关产品推荐

