You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于每行单元格的HTTP请求更新Pandas整列数据?

解决Pandas逐行调用API生成新列的问题

你当前代码的问题在于d['id']是Pandas的Series对象,直接拼进URL会把整列ID的字符串形式(比如类似0 123\n1 456\nName: id, dtype: int64)塞进请求地址,自然会触发InvalidURL错误。要实现每个ID单独请求API,得对id列的每个元素分别处理,下面是两种可行的实现方式:

基础实现(串行处理)

用apply函数对id列的每个值单独发起请求,适合数据量不大的场景:

import pandas as pd
import requests

def fetch_m0(id_val):
    try:
        # 用单个ID拼接URL
        resp = requests.get(f"http://localhost:3000/{id_val}")
        resp.raise_for_status()  # 捕获HTTP错误(比如404、500)
        # 解析JSON并提取目标字段
        return resp.json()['H']['M0']
    except Exception as err:
        # 打印错误信息,避免单个请求失败导致整个流程中断
        print(f"ID {id_val} 请求失败: {str(err)}")
        return None  # 出错时返回空值

# 假设你的DataFrame名为d,将结果赋值给m0列
d['m0'] = d['id'].apply(fetch_m0)

优化实现(并行处理)

如果你的ID数量较多,串行请求速度会很慢,可以用线程池并行发起请求,大幅提升效率:

import pandas as pd
import requests
from concurrent.futures import ThreadPoolExecutor

def fetch_m0(id_val):
    try:
        resp = requests.get(f"http://localhost:3000/{id_val}")
        resp.raise_for_status()
        return resp.json()['H']['M0']
    except Exception as err:
        print(f"ID {id_val} 请求失败: {str(err)}")
        return None

# 用线程池并行处理,max_workers可根据API并发限制调整
with ThreadPoolExecutor(max_workers=10) as executor:
    # 把id列转成列表,传给线程池处理
    d['m0'] = list(executor.map(fetch_m0, d['id'].tolist()))

注意事项

  • 确保已经安装requests库:执行pip install requests
  • 异常处理不可少,否则单个ID请求失败会导致整个列生成失败
  • 如果目标API有速率限制,要适当降低max_workers的值,或者在请求函数中加入time.sleep(0.1)这类延迟,避免被封禁

内容的提问来源于stack exchange,提问作者WoJ

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 18:10:34