You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas导出CSV仅保留最后一局数据问题排查求助

解决棒球赛事爬虫CSV导出仅保留最后一局数据的问题

爬取棒球赛事数据时,打印DataFrame能看到完整数据,但导出CSV后只剩最后一局的数据,相关代码如下:

#imports
from bs4 import BeautifulSoup
import requests
import pandas as pd
import csv

#Getting Website for each game ID in 2023 season
gameID = 598903
while gameID <= 598903: 
    website = f"https://baseball.pointstreak.com/boxscore.html?gameid={gameID}"
    result = requests.get(website)
    content = result.text
    gameID = gameID + 1
    #Convert html with Beatiful Soup
    soup = BeautifulSoup(content, 'lxml')
    #Get Inning playbyplay Data
    Innings = soup.findAll(class_='pbpinning')
    
#Strip text to clean data
for i in Innings:  
    #Cleaning \n and \t characterS
    data = ", ".join([stripped_strings for stripped_strings in i.stripped_strings])
    #print(data)
    
    #All data is pulled in as single string must structure it by inning then by play
    gamesum = data

    # Split the string into lines
    lines = gamesum.split('\n')

    # Iterate over each line and separate the inning and play
    for line in lines:
        parts = line.split(',', 1)
        inning = parts[0].strip()
        play = parts[1].strip()
        
        #Create a dataframe
        pbpdata = {'Inning': [inning], 'Play by Play': [play]}
        pbpdf = pd.DataFrame(pbpdata)
        print(pbpdf)
    
        #pbpdf.to_csv('baseball_game_play_by_play.csv', index = False)

错误原因

  1. 循环内重复创建DataFrame:最内层循环中每次都会新建pbpdf对象,之前循环生成的数据会被直接覆盖,最终仅保留最后一条记录。
  2. CSV写入模式默认覆盖:如果启用to_csv,默认使用mode='w'(覆盖模式),每次循环都会清空原有CSV文件,只写入当前行数据。

修正后的代码

from bs4 import BeautifulSoup
import requests
import pandas as pd

# 初始化存储所有赛事记录的空列表
all_plays = []

# 获取指定范围的比赛数据(此处为单场比赛)
gameID = 598903
while gameID <= 598903: 
    website = f"https://baseball.pointstreak.com/boxscore.html?gameid={gameID}"
    result = requests.get(website)
    content = result.text
    gameID += 1
    soup = BeautifulSoup(content, 'lxml')
    Innings = soup.findAll(class_='pbpinning')
    
    # 处理每一局的赛事数据
    for i in Innings:  
        # 清理特殊字符并拼接数据
        data = ", ".join([stripped_strings for stripped_strings in i.stripped_strings])
        lines = data.split('\n')
        
        # 拆分并提取每一条赛事记录
        for line in lines:
            parts = line.split(',', 1)
            # 增加格式校验,避免拆分失败导致索引错误
            if len(parts) == 2:
                inning = parts[0].strip()
                play = parts[1].strip()
                # 将单条记录追加到列表中
                all_plays.append({'Inning': inning, 'Play by Play': play})

# 所有记录收集完成后,统一转换为DataFrame
pbpdf = pd.DataFrame(all_plays)
print(pbpdf)

# 一次性导出完整数据到CSV文件
pbpdf.to_csv('baseball_game_play_by_play.csv', index=False)

关键修改说明

  • 用列表收集所有记录:初始化all_plays空列表,循环内将每条赛事记录以字典形式追加到列表,避免数据被覆盖。
  • 统一生成DataFrame:所有记录收集完成后再转换为DataFrame,保证数据完整性。
  • 单次写入CSV:仅执行一次to_csv操作,确保所有数据都写入文件,不会被覆盖。
  • 增加格式校验:添加if len(parts) == 2判断,避免因网页数据格式异常导致的程序报错。

内容的提问来源于stack exchange,提问作者Riley Bennett

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 07:58:41