You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

本地运行FastAPI时数据传输缓慢的原因排查

本地FastAPI调用耗时远超数据处理时间的原因及优化方案

核心问题:你的计时未覆盖序列化与响应发送环节

你代码里的print(time() - now)只统计了加载数据+生成字典的时间(7秒),但从json_data生成到把响应完整发送给请求方的过程完全没被计入——这部分才是耗时的核心,尤其是当数据集规模较大时。

具体原因及解决办法

1. pandas to_dict()生成的大字典导致JSON序列化极慢

如果你的数据集行数多、列数多,data.to_dict()会生成体积巨大的嵌套字典,FastAPI默认基于Python标准库json的序列化器处理这类大对象速度极慢。

  • 优化方案:
    • 直接用pandas的to_json()生成JSON字符串,跳过中间字典转换:
      @app.get('/data/{dataset}/{version}')
      def download_data(dataset: str, version: str):
          now = time()
          file = f"data/{dataset}/{dataset}_{version}.feather" 
          data = pd.read_feather(file)
          json_str = data.to_json(orient="records")
          print(time() - now)
          # 直接返回JSON字符串,FastAPI会自动处理响应格式
          return Response(content=json_str, media_type="application/json")
      
    • 替换FastAPI默认JSON编码器为更高效的orjson:
      先安装orjson,再初始化FastAPI时指定:
      from fastapi import FastAPI
      from fastapi.responses import ORJSONResponse
      
      app = FastAPI(default_response_class=ORJSONResponse)
      

2. 异步函数内跑同步阻塞代码拖慢事件循环

你用async def定义接口,但pd.read_feather()和data.to_dict()都是同步阻塞操作,会占用FastAPI的事件循环线程,哪怕单请求也会产生额外调度开销,降低处理效率。

  • 优化方案:
    • 把接口改成普通def(FastAPI会自动将同步函数放到线程池处理,反而更适配这类CPU密集型操作);
    • 或用asyncio.to_thread把同步操作放到单独线程:
      import asyncio
      
      @app.get('/data/{dataset}/{version}')
      async def download_data(dataset: str, version: str):
          now = time()
          def process_data():
              file = f"data/{dataset}/{dataset}_{version}.feather" 
              data = pd.read_feather(file)
              return data.to_dict()
          
          data_dict = await asyncio.to_thread(process_data)
          json_data = {"data": data_dict}
          print(time() - now)
          return json_data
      

3. 本地机器性能差异

序列化大JSON属于CPU密集型操作,若你的机器CPU核心数少、内存不足,耗时会远高于配置更好的同事机器——这也是你和同事耗时差距的主要原因之一。

4. 请求方的响应接收开销

requests.get()默认会把整个响应内容读到内存中,大JSON的解析和存储也会占用时间。可尝试流式读取减少本地开销:

r = requests.get('http://127.0.0.1:8000/data/dataset_name/version_name', stream=True)
# 逐块读取响应
for chunk in r.iter_content(chunk_size=8192):
    pass

验证方法

在代码中添加分段计时,明确各环节耗时:

@app.get('/data/{dataset}/{version}')
def download_data(dataset: str, version: str):
    now = time()
    file = f"data/{dataset}/{dataset}_{version}.feather" 
    data = pd.read_feather(file)
    print(f"数据加载耗时: {time() - now:.2f}s")
    
    now = time()
    data_dict = data.to_dict()
    print(f"字典生成耗时: {time() - now:.2f}s")
    
    now = time()
    import json
    json_str = json.dumps({"data": data_dict})
    print(f"JSON序列化耗时: {time() - now:.2f}s")
    
    return {"data": data_dict}

内容的提问来源于stack exchange,提问作者Juan C

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 16:03:22