You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何抓取对应日期的板球比赛数据?解决行数不匹配问题

解决板球赛事数据按日期关联的问题

核心思路是按日期区块遍历:网站中每个日期标题下方对应该日期的所有比赛,不能单独抓取所有比赛和所有日期,需要以日期为分组,逐个提取每个日期下的比赛数据,确保日期与比赛一一对应。

修改后的完整代码

import requests
import pandas as pd
from bs4 import BeautifulSoup

# 目标URL
url = 'https://halstead.play-cricket.com/Matches?fixture_month=13&home_or_away=both&page=1&q%5Bcategory_id%5D=all&q%5Bgender_id%5D=all&search_in=&season_id=255&seasonchange=f&selected_season_id=255&tab=Result&team_id=&utf8=%E2%9C%93&view_by=year'

data = requests.get(url).text
soup = BeautifulSoup(data, 'lxml')

# 准备存储最终数据的列表
final_data = []

# 处理第一个日期区块(无前置日期标题的情况)
first_date = soup.select_one('.col-sm-12.text-center.text-md-left.title2.padding_top_for_mobile')
if first_date:
    current_date = first_date.text.strip().replace('2022\n', '2022')
    # 获取当前日期下的所有比赛行
    next_date = first_date.find_next_sibling('.col-sm-12.text-center.text-md-left.title2.padding_top_for_mobile')
    if next_date:
        matches = first_date.find_next_siblings('.row.ml-large-0.mr-large-0', limit=next_date.index - first_date.index - 1)
    else:
        matches = first_date.find_next_siblings('.row.ml-large-0.mr-large-0')
    
    # 解析每场比赛信息
    for match in matches:
        teams = match.select('p.txt1')
        team_a = teams[0].text.strip() if len(teams)>=2 else 'N/A'
        team_b = teams[1].text.strip() if len(teams)>=2 else 'N/A'
        result = match.select_one('.fonts-gt').text.strip().replace('  ', ' ') if match.select_one('.fonts-gt') else 'N/A'
        
        final_data.append({
            'Date': current_date,
            'Team A': team_a,
            'Team B': team_b,
            'Result': result
        })

# 处理后续日期区块
date_titles = soup.select('.col-sm-12.text-center.text-md-left.title2.padding_top_for_mobile')[1:]
for title in date_titles:
    current_date = title.text.strip().replace('2022\n', '2022')
    # 获取当前日期对应的比赛行
    match_row = title.find_next_sibling('.row.ml-large-0.mr-large-0')
    if match_row:
        teams = match_row.select('p.txt1')
        team_a = teams[0].text.strip() if len(teams)>=2 else 'N/A'
        team_b = teams[1].text.strip() if len(teams)>=2 else 'N/A'
        result = match_row.select_one('.fonts-gt').text.strip().replace('  ', ' ') if match_row.select_one('.fonts-gt') else 'N/A'
        
        final_data.append({
            'Date': current_date,
            'Team A': team_a,
            'Team B': team_b,
            'Result': result
        })

# 转换为DataFrame并输出
df = pd.DataFrame(final_data)
print(f"最终DataFrame行数:{len(df)}")
print(df)

关键逻辑说明

  • 按日期分组绑定:以日期标题为锚点,提取其下方对应的所有比赛行,确保每场比赛都关联到正确的日期,从根源避免行数不匹配问题。
  • 边界情况处理:单独处理首个日期区块(无前置标题的特殊场景),同时对比赛信息缺失的情况做兜底赋值。
  • 结构化存储:用字典列表统一存储单条比赛数据,最后直接转换为DataFrame,保证数据结构一致性。

内容的提问来源于stack exchange,提问作者Mark W

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 01:55:16