You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页抓取:多URL循环爬取球员数据的代码问题排查

代码错误分析与修正方案

原代码存在的核心问题

  • 函数定义位置错误:get_data函数被嵌套在URL循环内部,既不符合代码规范,也会导致重复定义逻辑混乱。
  • 请求参数错误:requests.get(urllist)传入的是URL列表而非当前循环的单个URL,会直接触发请求异常。
  • 节点遍历错误:直接遍历find获取的表格对象,会遍历表格的所有子节点(包括表头、空标签等),而非球员数据行,应提取表格内的<tr>元素。
  • 未定义变量引用:代码中使用了未初始化的data.append(item),实际应使用已声明的playerdata。
  • 未执行爬取函数:仅定义了get_data但从未调用,导致没有任何数据被爬取。
  • 未区分数据类型:没有根据URL后缀(ballkontakte/fairplay)针对性提取数据,会导致两个URL的数据混在一起,出现大量无效空值。

修正后的完整代码

from bs4 import BeautifulSoup
import requests
import pandas as pd

# 生成目标URL列表
changing_words = ["-ballkontakte/", "-fairplay/"]
root_url = "https://sportdaten.spiegel.de/fussball/bundesliga/ma9417803/fc-augsburg_eintracht-frankfurt/spielstatistik"
url_list = [root_url + word for word in changing_words]

# 用字典存储球员数据,以姓名为键方便合并不同URL的同球员信息
player_dict = {}

def scrape_player_data(url):
    response = requests.get(url)
    response.raise_for_status()  # 检查请求是否成功,失败则抛出异常
    soup = BeautifulSoup(response.content, "lxml")
    
    # 获取表格并提取数据行(跳过表头行)
    table = soup.find("table", class_="module-statistics statistics")
    if not table:
        return
    rows = table.find_all("tr")[1:]  # 从第2行开始,跳过表头
    
    # 根据URL后缀判断要提取的数据类型
    is_ball_data = "-ballkontakte/" in url
    is_fairplay_data = "-fairplay/" in url
    
    for row in rows:
        name_td = row.find("td", class_="person-name")
        if not name_td:
            continue
        player_name = name_td.text.strip()
        
        # 初始化球员数据(若不存在于字典中)
        if player_name not in player_dict:
            player_dict[player_name] = {
                "Name": player_name,
                "Balls touched": "",
                "Yellow Cards": ""
            }
        
        # 针对性提取对应数据
        if is_ball_data:
            balls_td = row.find("td", class_="person_stats-balls_touched person_stats-balls_touched-list")
            if balls_td:
                player_dict[player_name]["Balls touched"] = balls_td.text.strip()
        elif is_fairplay_data:
            yellow_td = row.find("td", class_="person_stats-card_yellow person_stats-card_yellow-list")
            if yellow_td:
                player_dict[player_name]["Yellow Cards"] = yellow_td.text.strip()

# 遍历所有URL执行爬取
for url in url_list:
    scrape_player_data(url)

# 转换为列表格式便于后续处理
player_data = list(player_dict.values())

# 生成需求的两个目标列表
# 1. 所有球员姓名及触球数(过滤无数据条目)
players_with_balls = [
    {"Name": p["Name"], "Balls touched": p["Balls touched"]} 
    for p in player_data if p["Balls touched"]
]
# 2. 有黄牌的球员及其数量(过滤无黄牌条目)
players_with_yellow = [
    {"Name": p["Name"], "Yellow Cards": p["Yellow Cards"]} 
    for p in player_data if p["Yellow Cards"]
]

# 输出结果
print("所有球员姓名及触球数:")
for item in players_with_balls:
    print(item)

print("\n有黄牌的球员及其数量:")
for item in players_with_yellow:
    print(item)

# 可选:将完整数据转为DataFrame保存或查看
df = pd.DataFrame(player_data)
print("\n完整球员数据表格:")
print(df)

修正说明

  1. 把爬取函数移到循环外部,规范代码结构,避免重复定义。
  2. 使用单个URL发起请求,添加请求状态检查,便于排查网络问题。
  3. 提取表格的<tr>数据行并跳过表头,确保只处理球员数据。
  4. 用字典存储球员数据,通过姓名作为唯一键,自动合并不同URL的同球员信息,避免重复条目。
  5. 根据URL后缀判断数据类型,针对性提取触球数或黄牌数,避免无效空值。
  6. 明确生成需求的两个目标列表,过滤掉无对应数据的条目。

内容的提问来源于stack exchange,提问作者Dominik Kacinski

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 05:22:50