You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

加载CSV球员ID传入URL参数爬取NBA数据失败如何解决

NBA球员数据爬虫问题修复方案

核心问题点

  • 接口URL格式错误:使用三引号换行定义的URL携带了大量空白字符、冗余参数和写死的PlayerID,和后续传入的params参数冲突,导致请求参数解析异常
  • 协议使用错误:stats.nba.com接口现在仅支持HTTPS访问,使用HTTP会导致请求被拒绝或跳转
  • 读取的球员ID未做清洗:从CSV读取的ID可能携带空格、换行等不可见字符,和接口要求的参数格式不匹配
  • 参数编码错误:params中的SeasonType手动写了+号,requests会自动对空格做URL编码,手动加+会导致参数被二次编码,接口识别失败
  • 数据处理逻辑缺失:生成DataFrame时未传入抓取到的rowSet数据,也未将单球员数据合并到全局存储的best_db中,df.head是属性,需要调用df.head()才能输出预览内容

修复后代码

import pandas as pd
import requests
import csv

best_db = pd.DataFrame()

def table_Scrape():
    global best_db 
    
    with open("SHORT_ID_plyr.csv", "r", encoding="utf-8") as f_urls: 
        f_urls_list = csv.reader(f_urls, delimiter=',') 
        next(f_urls_list)    
        
        # 定义干净的接口根地址,不要携带任何参数和空白字符
        base_url = "https://stats.nba.com/stats/playergamelogs"
        header_dict = {
            'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/110.0.0.0 Safari/537.36',
            'x-nba-stats-origin': 'stats',
            'x-nba-stats-token': 'true',
            'Referer': 'https://stats.nba.com',
            'Connection': 'keep-alive'
        }

        for lines in f_urls_list:        
            # 清洗球员ID,去除空白字符
            player_id = lines[0].strip()
            print(f"正在抓取球员ID: {player_id}")

            params = {
                'LastNGames': '0',
                'LeagueID': '00',
                'MeasureType': 'Base',
                'Month': '0',
                'OpponentTeamID': '0',
                'PORound': '0',
                'PaceAdjust': 'N',
                'PerMode': 'Totals',
                'Period': '0',
                'PlayerID': player_id,
                'PlusMinus': 'N',
                'Rank': 'N',
                'Season': '2021-22',
                'SeasonType': 'Regular Season',
                'DateFrom':'',
                'DateTo':'',
                'GameSegment':'',
                'Location':'',
                'Outcome':'',
                'SeasonSegment':'',
                'ShotClockRange':'',
                'VsConference':'',
                'VsDivision':''
            }

            # 发送请求,添加异常判断避免单次抓取失败导致程序终止
            try:
                res = requests.get(base_url, headers=header_dict, params=params, timeout=10)
                res.raise_for_status()
                json_set = res.json()
                headers = json_set['resultSets'][0]['headers']
                data_set = json_set['resultSets'][0]['rowSet']
                # 生成单球员数据DataFrame并合并到全局数据集
                df = pd.DataFrame(data_set, columns=headers)
                best_db = pd.concat([best_db, df], ignore_index=True)
                print(f"球员ID {player_id} 抓取完成,累计数据条数:{len(best_db)}")
                # 如需预览单球员数据可调用 print(df.head())
            except Exception as e:
                print(f"球员ID {player_id} 抓取失败,错误信息:{str(e)}")
                continue

table_Scrape()
# 全部抓取完成后可执行下方代码导出结果
# best_db.to_csv("player_game_logs_2021-22.csv", index=False, encoding="utf-8-sig")

内容的提问来源于stack exchange,提问作者trambilo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 18:54:04