You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup抓取同一div下不同data-stat属性的数据?

Basketball Reference爬虫批量抓取data-stat数据问题解决

问题描述

我编写了基于BeautifulSoup的basketball-reference.com爬虫,目前仅能通过变量赋值或用户输入获取单个data-stat属性(如pg_per_mp)对应的数据。希望定义一个data-stat列表,批量抓取列表中每个属性对应的数据,但修改代码后未成功。

原代码

from bs4 import BeautifulSoup
import requests

first = ()
first_slice = ()
last = ()


def askname():
    global first
    first = input(str("First Name of Player?"))
    global last
    last = input(str("Last Name of Player?"))
    print("Confirmed, loading up " + first + " " + last)
# asks user for player name

askname()

first_slice_result = (first[:2])
last_slice_result = (last[:5])
print(first_slice_result)
print(last_slice_result)
# slices player's name so it can match the format bref uses
first_slice_resultA = str(first_slice_result)
last_slice_resultA = str(last_slice_result)

first_last_slice = last_slice_resultA + first_slice_resultA

lower = first_last_slice.lower() + "01"

start_letter = (last[:1])
lower_letter = (start_letter.lower())
# grabs the letter bref uses for organization

print(lower)
source = requests.get('https://www.basketball-reference.com/players/' + lower_letter + '/' + lower + '.html').text

soup = BeautifulSoup(source, 'lxml')
tbody = soup.find('tbody')
pergame = tbody.find(class_="full_table")
classrite = tbody.find(class_="right")
tr_body = tbody.find_all('tr')
# lprint(pergame)

for td in tbody:
    print(td.get_text)

print("done")

get = str(input("What stat? \nCheck commands.txt for statistic names. \n"))

for trb in tr_body:
    print(trb.get('id'))
    print("\n")

    th = trb.find('th')
    print(th.get_text())
    print(th.get('data-stat'))

    row = {}
    for td in trb.find_all('td'):
        row[td.get('data-stat')] = td.get_text()

    print(row[get])

错误的修改尝试

get = [fga_per_mp, fg3_per_mp]

for trb in tr_body:
    print(trb.get('id'))
    print("\n")

    th = trb.find('th')
    print(th.get_text())
    print(th.get('data-stat'))

    row = {}
    for td in trb.find_all('td'):
        for x in get():
          row[td.get('data-stat')] = td.get_text()

问题分析与解决方法

错误点说明

  1. 列表元素未加引号:[fga_per_mp, fg3_per_mp]中的元素会被识别为变量而非字符串,导致未定义错误,需改为字符串列表['fga_per_mp', 'fg3_per_mp']。
  2. 列表循环语法错误:for x in get()是错误调用,get是列表,直接用for x in get即可,无需加括号。
  3. 逻辑错误:嵌套循环会重复覆盖row字典,正确逻辑应先把当前行所有data-stat数据存入字典,再遍历目标列表提取对应值。

正确修改代码

替换原代码中用户输入get及后续循环部分,改为批量抓取逻辑:

# 定义需要批量抓取的data-stat属性列表
target_stats = ['fga_per_mp', 'fg3_per_mp', 'ft_per_mp']

for trb in tr_body:
    # 跳过空行(部分行是分隔线,无数据)
    if trb.get('class') and 'thead' in trb.get('class'):
        continue
    
    print(trb.get('id'))
    print("\n")

    th = trb.find('th')
    season_name = th.get_text()
    print(f"赛季: {season_name}")

    row = {}
    # 先将当前行所有data-stat数据存入字典
    for td in trb.find_all('td'):
        row[td.get('data-stat')] = td.get_text()
    
    # 批量输出目标属性的数据
    print("目标数据:")
    for stat in target_stats:
        # 使用get方法避免属性不存在时抛出KeyError
        stat_value = row.get(stat, "无数据")
        print(f"- {stat}: {stat_value}")
    print("-" * 30)

额外优化建议

  • 删除原代码中无效的for td in tbody循环,该循环会打印tbody的所有子节点(包括tr),输出冗余内容。
  • 可以去掉全局变量,将askname函数改为返回姓名元组,更符合Python规范。

内容的提问来源于stack exchange,提问作者Xendex

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 19:05:47