You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

解决API返回CSV格式数据无法加载为Pandas DataFrame的问题

解决Pandas读取GENESIS API响应的ParserError问题

问题分析

报错原因是API响应混杂了大量描述性文本,数据采用分号(;)而非默认逗号分隔,且存在大量多余连续分号,导致Pandas无法正确识别字段数量。

解决步骤

以下是具体修改方案:

  1. 过滤无效行并清洗数据
    先拆分响应文本,过滤开头的说明内容,清理每行多余的连续分号,合并类别与对应数据行:
import requests
import pandas as pd
import io
import re

# 获取API响应(保留你的原有代码)
response = requests.request("GET", url, headers=headers, data=payload)

# 拆分响应为行列表
lines = response.text.splitlines()

# 定位数据起始行(找到包含月份表头的行)
start_idx = None
for i, line in enumerate(lines):
    if "January" in line:
        start_idx = i
        break

# 处理数据行:清理多余分号,合并类别与数据
processed_lines = []
current_category = ""
for line in lines[start_idx:]:
    # 替换多个连续分号为单个,去除首尾分号
    cleaned_line = re.sub(r";+", ";", line).strip(";")
    if not cleaned_line:
        continue
    # 判断是否为类别行(无数字,仅分类名称)
    if not any(c.isdigit() for c in cleaned_line):
        current_category = cleaned_line
    else:
        # 将类别作为第一列拼接数据行
        processed_lines.append(f"{current_category};{cleaned_line}")

# 手动构造表头
headers = [
    "Category", "Price Type", "Adjustment Type",
    "January", "February", "March", "April", "May", "June",
    "July", "August", "September", "October", "November", "December"
]
  1. 读取清洗后的数据为DataFrame
# 用io.StringIO包装处理后的文本,指定分号分隔符
df = pd.read_csv(
    io.StringIO("\n".join(processed_lines)),
    sep=";",
    header=None,
    names=headers
)

额外处理建议

  • 若响应中存在...形式的缺失值,可执行df.replace("...", pd.NA)转为标准缺失值格式,方便后续分析。
  • 若类别行与数据行的对应规则有变动,可调整类别行的判断逻辑(比如根据行内分号数量区分)。

内容的提问来源于stack exchange,提问作者prashanth manohar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 14:43:26