本地运行FastAPI时数据传输缓慢的原因排查
本地FastAPI调用耗时远超数据处理时间的原因及优化方案
核心问题:你的计时未覆盖序列化与响应发送环节
你代码里的print(time() - now)只统计了加载数据+生成字典的时间(7秒),但从json_data生成到把响应完整发送给请求方的过程完全没被计入——这部分才是耗时的核心,尤其是当数据集规模较大时。
具体原因及解决办法
1. pandas to_dict()生成的大字典导致JSON序列化极慢
如果你的数据集行数多、列数多,data.to_dict()会生成体积巨大的嵌套字典,FastAPI默认基于Python标准库json的序列化器处理这类大对象速度极慢。
- 优化方案:
- 直接用pandas的
to_json()生成JSON字符串,跳过中间字典转换:@app.get('/data/{dataset}/{version}') def download_data(dataset: str, version: str): now = time() file = f"data/{dataset}/{dataset}_{version}.feather" data = pd.read_feather(file) json_str = data.to_json(orient="records") print(time() - now) # 直接返回JSON字符串,FastAPI会自动处理响应格式 return Response(content=json_str, media_type="application/json") - 替换FastAPI默认JSON编码器为更高效的
orjson:
先安装orjson,再初始化FastAPI时指定:from fastapi import FastAPI from fastapi.responses import ORJSONResponse app = FastAPI(default_response_class=ORJSONResponse)
- 直接用pandas的
2. 异步函数内跑同步阻塞代码拖慢事件循环
你用async def定义接口,但pd.read_feather()和data.to_dict()都是同步阻塞操作,会占用FastAPI的事件循环线程,哪怕单请求也会产生额外调度开销,降低处理效率。
- 优化方案:
- 把接口改成普通
def(FastAPI会自动将同步函数放到线程池处理,反而更适配这类CPU密集型操作); - 或用
asyncio.to_thread把同步操作放到单独线程:import asyncio @app.get('/data/{dataset}/{version}') async def download_data(dataset: str, version: str): now = time() def process_data(): file = f"data/{dataset}/{dataset}_{version}.feather" data = pd.read_feather(file) return data.to_dict() data_dict = await asyncio.to_thread(process_data) json_data = {"data": data_dict} print(time() - now) return json_data
- 把接口改成普通
3. 本地机器性能差异
序列化大JSON属于CPU密集型操作,若你的机器CPU核心数少、内存不足,耗时会远高于配置更好的同事机器——这也是你和同事耗时差距的主要原因之一。
4. 请求方的响应接收开销
requests.get()默认会把整个响应内容读到内存中,大JSON的解析和存储也会占用时间。可尝试流式读取减少本地开销:
r = requests.get('http://127.0.0.1:8000/data/dataset_name/version_name', stream=True) # 逐块读取响应 for chunk in r.iter_content(chunk_size=8192): pass
验证方法
在代码中添加分段计时,明确各环节耗时:
@app.get('/data/{dataset}/{version}') def download_data(dataset: str, version: str): now = time() file = f"data/{dataset}/{dataset}_{version}.feather" data = pd.read_feather(file) print(f"数据加载耗时: {time() - now:.2f}s") now = time() data_dict = data.to_dict() print(f"字典生成耗时: {time() - now:.2f}s") now = time() import json json_str = json.dumps({"data": data_dict}) print(f"JSON序列化耗时: {time() - now:.2f}s") return {"data": data_dict}
内容的提问来源于stack exchange,提问作者Juan C
相关产品推荐
相关产品推荐

