You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas爬取NCAA数据时Score字段格式异常问题求助

NCAA赛事数据爬取:Score字段格式异常解决方法
import requests
import pandas as pd
import csv


headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/103.0.0.0 Safari/537.36"
}

urls = [
    'https://stats.ncaa.org/game/play_by_play/12465',
    'https://stats.ncaa.org/game/play_by_play/12755',
    'https://stats.ncaa.org/game/play_by_play/12640',
    'https://stats.ncaa.org/game/play_by_play/12290',
]
s = requests.Session()
s.headers.update(headers)
for url in urls:
    r = s.get(url)
    dfs = pd.read_html(r.text)
    for df in dfs:
        df.to_csv('pbp.csv', mode='a', index=False)

我借助社区帮助完成了这段Python脚本,可爬取上述多个NCAA赛事逐场数据页面的表格内容并导出至CSV文件。但部分Score字段出现格式异常(如下图所示),尝试在Excel/Sheets中格式化也无法解决。检查页面元素未发现这些异常条目有特殊之处,仅每个链接中存在少量此类情况,请问如何解决该问题?

格式异常的Score字段


问题原因

这种异常是页面中存在不可见特殊字符(比如Unicode控制字符、全角空格、零宽字符等)导致的,pd.read_html会直接读取这些字符,导出CSV后就会显示乱码或异常格式。

解决方法

在导出CSV前,对Score字段做清洗,过滤掉非必要字符,只保留数字和比分分隔符(比如-):

import requests
import pandas as pd
import re

headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/103.0.0.0 Safari/537.36"
}

urls = [
    'https://stats.ncaa.org/game/play_by_play/12465',
    'https://stats.ncaa.org/game/play_by_play/12755',
    'https://stats.ncaa.org/game/play_by_play/12640',
    'https://stats.ncaa.org/game/play_by_play/12290',
]

# 定义清洗函数:仅保留数字和连字符
def clean_score(score_str):
    if pd.isna(score_str):
        return score_str
    # 过滤所有非数字、非连字符的字符
    cleaned = re.sub(r'[^0-9-]', '', str(score_str))
    return cleaned if cleaned else score_str

s = requests.Session()
s.headers.update(headers)
first_write = True  # 避免重复写入表头

for url in urls:
    r = s.get(url)
    dfs = pd.read_html(r.text)
    for df in dfs:
        # 检查并清洗Score列
        if 'Score' in df.columns:
            df['Score'] = df['Score'].apply(clean_score)
        # 控制表头仅写入一次
        df.to_csv('pbp.csv', mode='a', index=False, header=first_write)
        first_write = False

说明

  1. 清洗函数用正则表达式过滤掉所有非数字和连字符的内容,确保Score字段仅保留比分的核心信息
  2. 添加first_write标志,解决原脚本重复写入表头的问题
  3. 运行修改后的脚本后,导出的CSV中Score字段格式会恢复正常,Excel/Sheets可以正常识别

内容的提问来源于stack exchange,提问作者Anthony Madle

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 15:31:05