如何实现read_html循环?Python多HLTV赛事数据爬取问题
解决HLTV赛事数据爬取:获取当日所有赛事并遍历整页
我明白你的问题了——当前代码只能抓取每个赛事组里的第一场比赛数据,重复填充的问题是因为你没有正确遍历每场赛事的链接和对应信息。咱们一步步调整代码,实现整页所有赛事的爬取:
核心问题分析
- 原代码里
link = table.find('a', href=True)只取了第一个赛事链接,但每个results-sublist(当日赛事组)里可能有多场比赛,需要遍历所有a标签。 - 你直接用
df2[0].iloc[0]取首场比赛的队伍信息,没有遍历df2[0]的每一行,导致所有赛事都复用了首场的数据。 - 处理选图信息的逻辑需要嵌套到每场赛事的循环里,确保每场比赛的选图字典都是独立初始化的,避免数据残留。
修改后的完整代码
import requests from bs4 import BeautifulSoup as bs import pandas as pd import re from tabulate import tabulate team_id = 8362 team_name = "your_team_name" # 记得替换成实际队伍名称 list_dfs = [] # 遍历分页(这里示例是前1页,你可以调整range参数爬更多页) for page_offset in range(0, 1): offset = page_offset * 100 url = f"https://www.hltv.org/results?offset={offset}&team={team_id}" res = requests.get(url) soup = bs(res.content, 'lxml') tables = soup.find_all("div", {"class": "results-sublist"}) # 遍历每个当日赛事组 for table in tables: # 获取该组所有赛事的表格数据 df2 = pd.read_html(str(table))[0] # 获取该组所有赛事的详情链接 links = table.find_all('a', href=True) # 遍历每场赛事:同时匹配表格行和对应链接 for idx, (row, link) in enumerate(zip(df2.itertuples(), links)): # 初始化单场比赛的选图信息字典 dict_choices = {"teamchoose": [], "chosen": [], "maps": []} # 构造完整的赛事详情链接 match_link = f"https://www.hltv.org/{link.get('href')}" match_res = requests.get(match_link) match_soup = bs(match_res.content, 'lxml') # 解析比赛日期 unix_time = int(match_soup.select(".timeAndEvent div")[0]['data-unix']) date = pd.to_datetime(unix_time * 1000000) # 提取选图步骤信息 temp = match_soup.find_all("div", {"class": "padding"}) out = re.findall(r'<div>\d\.(.*?)</div>', str(temp)) # 处理常规ban/pick步骤 for choice in out[:6]: split = choice.strip().split(" ") dict_choices["teamchoose"].append(" ".join(split[:-2])) dict_choices["chosen"].append(split[-2]) dict_choices["maps"].append(split[-1]) # 处理可能存在的决胜图 try: left = out[6] split = left.strip().split(" ") dict_choices["teamchoose"].append(split[2]) dict_choices["chosen"].append(split[2]) dict_choices["maps"].append(split[0]) except IndexError: pass # 无决胜图时跳过 # 构造单场比赛的DataFrame match_df = pd.DataFrame.from_dict(dict_choices, orient='index').transpose() match_df["opponent"] = row[2] # 从表格行获取对手信息 match_df["team"] = row[0] # 从表格行获取当前队伍信息 match_df["match"] = len(list_dfs) # 全局唯一赛事索引 match_df['date'] = date # 将单场数据加入全局列表 list_dfs.append(match_df) # 合并所有数据并整理列顺序 df_out = pd.concat(list_dfs, ignore_index=True) df_out = df_out[['match','date','team','opponent','teamchoose','chosen','maps']] # 保存到CSV并打印结果 df_out.to_csv(f"{team_name}_vetoes.csv", index=False) print(tabulate(df_out, headers='keys', tablefmt='psql'))
关键修改点说明
- 遍历所有赛事链接:用
table.find_all('a', href=True)获取该组所有赛事的详情链接,不再只取第一个。 - 逐行匹配赛事数据:通过
zip(df2.itertuples(), links)同时遍历表格行和对应链接,确保每场比赛的队伍信息和详情页一一对应。 - 独立初始化选图字典:把
dict_choices的初始化放到每场赛事的循环内,彻底避免之前比赛的数据残留。 - 全局赛事索引:用
len(list_dfs)作为match字段的值,保证每个赛事有唯一的标识。 - 分页逻辑优化:将
list_dfs移到外层循环外,避免每页数据被覆盖丢失。
这样修改后,就能完整爬取整页所有赛事的数据,不会再出现重复填充的问题啦!
内容的提问来源于stack exchange,提问作者Kreeshee
相关产品推荐
相关产品推荐

