You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何实现read_html循环?Python多HLTV赛事数据爬取问题

解决HLTV赛事数据爬取:获取当日所有赛事并遍历整页

我明白你的问题了——当前代码只能抓取每个赛事组里的第一场比赛数据,重复填充的问题是因为你没有正确遍历每场赛事的链接和对应信息。咱们一步步调整代码,实现整页所有赛事的爬取:

核心问题分析

  1. 原代码里link = table.find('a', href=True)只取了第一个赛事链接,但每个results-sublist(当日赛事组)里可能有多场比赛,需要遍历所有a标签。
  2. 你直接用df2[0].iloc[0]取首场比赛的队伍信息,没有遍历df2[0]的每一行,导致所有赛事都复用了首场的数据。
  3. 处理选图信息的逻辑需要嵌套到每场赛事的循环里,确保每场比赛的选图字典都是独立初始化的,避免数据残留。

修改后的完整代码

import requests
from bs4 import BeautifulSoup as bs
import pandas as pd
import re
from tabulate import tabulate

team_id = 8362
team_name = "your_team_name"  # 记得替换成实际队伍名称
list_dfs = []

# 遍历分页(这里示例是前1页,你可以调整range参数爬更多页)
for page_offset in range(0, 1):
    offset = page_offset * 100
    url = f"https://www.hltv.org/results?offset={offset}&team={team_id}"
    res = requests.get(url)
    soup = bs(res.content, 'lxml')
    tables = soup.find_all("div", {"class": "results-sublist"})
    
    # 遍历每个当日赛事组
    for table in tables:
        # 获取该组所有赛事的表格数据
        df2 = pd.read_html(str(table))[0]
        # 获取该组所有赛事的详情链接
        links = table.find_all('a', href=True)
        
        # 遍历每场赛事:同时匹配表格行和对应链接
        for idx, (row, link) in enumerate(zip(df2.itertuples(), links)):
            # 初始化单场比赛的选图信息字典
            dict_choices = {"teamchoose": [], "chosen": [], "maps": []}
            # 构造完整的赛事详情链接
            match_link = f"https://www.hltv.org/{link.get('href')}"
            match_res = requests.get(match_link)
            match_soup = bs(match_res.content, 'lxml')
            
            # 解析比赛日期
            unix_time = int(match_soup.select(".timeAndEvent div")[0]['data-unix'])
            date = pd.to_datetime(unix_time * 1000000)
            
            # 提取选图步骤信息
            temp = match_soup.find_all("div", {"class": "padding"})
            out = re.findall(r'<div>\d\.(.*?)</div>', str(temp))
            
            # 处理常规ban/pick步骤
            for choice in out[:6]:
                split = choice.strip().split(" ")
                dict_choices["teamchoose"].append(" ".join(split[:-2]))
                dict_choices["chosen"].append(split[-2])
                dict_choices["maps"].append(split[-1])
            
            # 处理可能存在的决胜图
            try:
                left = out[6]
                split = left.strip().split(" ")
                dict_choices["teamchoose"].append(split[2])
                dict_choices["chosen"].append(split[2])
                dict_choices["maps"].append(split[0])
            except IndexError:
                pass  # 无决胜图时跳过
            
            # 构造单场比赛的DataFrame
            match_df = pd.DataFrame.from_dict(dict_choices, orient='index').transpose()
            match_df["opponent"] = row[2]  # 从表格行获取对手信息
            match_df["team"] = row[0]      # 从表格行获取当前队伍信息
            match_df["match"] = len(list_dfs)  # 全局唯一赛事索引
            match_df['date'] = date
            
            # 将单场数据加入全局列表
            list_dfs.append(match_df)

# 合并所有数据并整理列顺序
df_out = pd.concat(list_dfs, ignore_index=True)
df_out = df_out[['match','date','team','opponent','teamchoose','chosen','maps']]

# 保存到CSV并打印结果
df_out.to_csv(f"{team_name}_vetoes.csv", index=False)
print(tabulate(df_out, headers='keys', tablefmt='psql'))

关键修改点说明

  • 遍历所有赛事链接:用table.find_all('a', href=True)获取该组所有赛事的详情链接,不再只取第一个。
  • 逐行匹配赛事数据:通过zip(df2.itertuples(), links)同时遍历表格行和对应链接,确保每场比赛的队伍信息和详情页一一对应。
  • 独立初始化选图字典:把dict_choices的初始化放到每场赛事的循环内,彻底避免之前比赛的数据残留。
  • 全局赛事索引:用len(list_dfs)作为match字段的值,保证每个赛事有唯一的标识。
  • 分页逻辑优化:将list_dfs移到外层循环外,避免每页数据被覆盖丢失。

这样修改后,就能完整爬取整页所有赛事的数据,不会再出现重复填充的问题啦!

内容的提问来源于stack exchange,提问作者Kreeshee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 13:32:28