如何抓取对应日期的板球比赛数据?解决行数不匹配问题
解决板球赛事数据按日期关联的问题
核心思路是按日期区块遍历:网站中每个日期标题下方对应该日期的所有比赛,不能单独抓取所有比赛和所有日期,需要以日期为分组,逐个提取每个日期下的比赛数据,确保日期与比赛一一对应。
修改后的完整代码
import requests import pandas as pd from bs4 import BeautifulSoup # 目标URL url = 'https://halstead.play-cricket.com/Matches?fixture_month=13&home_or_away=both&page=1&q%5Bcategory_id%5D=all&q%5Bgender_id%5D=all&search_in=&season_id=255&seasonchange=f&selected_season_id=255&tab=Result&team_id=&utf8=%E2%9C%93&view_by=year' data = requests.get(url).text soup = BeautifulSoup(data, 'lxml') # 准备存储最终数据的列表 final_data = [] # 处理第一个日期区块(无前置日期标题的情况) first_date = soup.select_one('.col-sm-12.text-center.text-md-left.title2.padding_top_for_mobile') if first_date: current_date = first_date.text.strip().replace('2022\n', '2022') # 获取当前日期下的所有比赛行 next_date = first_date.find_next_sibling('.col-sm-12.text-center.text-md-left.title2.padding_top_for_mobile') if next_date: matches = first_date.find_next_siblings('.row.ml-large-0.mr-large-0', limit=next_date.index - first_date.index - 1) else: matches = first_date.find_next_siblings('.row.ml-large-0.mr-large-0') # 解析每场比赛信息 for match in matches: teams = match.select('p.txt1') team_a = teams[0].text.strip() if len(teams)>=2 else 'N/A' team_b = teams[1].text.strip() if len(teams)>=2 else 'N/A' result = match.select_one('.fonts-gt').text.strip().replace(' ', ' ') if match.select_one('.fonts-gt') else 'N/A' final_data.append({ 'Date': current_date, 'Team A': team_a, 'Team B': team_b, 'Result': result }) # 处理后续日期区块 date_titles = soup.select('.col-sm-12.text-center.text-md-left.title2.padding_top_for_mobile')[1:] for title in date_titles: current_date = title.text.strip().replace('2022\n', '2022') # 获取当前日期对应的比赛行 match_row = title.find_next_sibling('.row.ml-large-0.mr-large-0') if match_row: teams = match_row.select('p.txt1') team_a = teams[0].text.strip() if len(teams)>=2 else 'N/A' team_b = teams[1].text.strip() if len(teams)>=2 else 'N/A' result = match_row.select_one('.fonts-gt').text.strip().replace(' ', ' ') if match_row.select_one('.fonts-gt') else 'N/A' final_data.append({ 'Date': current_date, 'Team A': team_a, 'Team B': team_b, 'Result': result }) # 转换为DataFrame并输出 df = pd.DataFrame(final_data) print(f"最终DataFrame行数:{len(df)}") print(df)
关键逻辑说明
- 按日期分组绑定:以日期标题为锚点,提取其下方对应的所有比赛行,确保每场比赛都关联到正确的日期,从根源避免行数不匹配问题。
- 边界情况处理:单独处理首个日期区块(无前置标题的特殊场景),同时对比赛信息缺失的情况做兜底赋值。
- 结构化存储:用字典列表统一存储单条比赛数据,最后直接转换为DataFrame,保证数据结构一致性。
内容的提问来源于stack exchange,提问作者Mark W
相关产品推荐
相关产品推荐

