You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python大字典快速迭代优化:提升大JSON情感分析运行速度

大型JSON情感分析任务速度优化方案

你当前代码运行慢的核心原因是串行执行HTTP请求:循环里每发一次API请求都要等到返回结果才会处理下一条,绝大多数耗时都浪费在网络IO等待上,普通的map()函数本质还是同步串行执行,完全解决不了IO阻塞的问题,不适合这个场景。

为什么不推荐用map()

  • 内置map()是同步迭代逻辑,和你现在写的for循环执行效率没有本质差异,还是会一条一条等接口返回,无法缩短总耗时
  • map()的返回值是迭代器,你需要额外遍历赋值才能修改原始字典的内容,代码反而更冗余

可落地的优化方案

1. 核心优化:多线程并发请求

网络IO场景下用多线程并发可以同时发起几十条请求,总耗时会直接降到原来串行的1/N(N为并发数),是收益最高的优化手段。用Python内置的concurrent.futures模块即可实现,不需要引入复杂的第三方异步框架。
优化后的完整实现参考:

import json
import zipfile
import requests
from concurrent.futures import ThreadPoolExecutor, as_completed

# 全局复用Session,自动复用TCP连接,减少握手开销
SESSION = requests.Session()
# 支持的语言存为集合,in判断速度远快于列表
SUPPORTED_LANG = {"en", "es", "ar", "da", "de", "fr", "it", "ja", "nl", "pl", "pt", "ru", "sv", "sw", "zh"}

class Pipeline:

    def __init__(self, json_file_path, json_file_path_zip=None, API="--"):
        '''Initializes a Pipeline object.
        INPUT:
        json_file_path: A json file path
        API: the current API'''

        self.API = API
        self.json_file_path = json_file_path
        self.json_file_path_zip = json_file_path + '.zip'
        self.json_file = self.converter()

    
    def converter(self):
        with zipfile.ZipFile(self.json_file_path_zip, 'r') as zip_ref:
            zip_ref.extractall()

        with open(self.json_file_path) as ff:
            d = json.load(ff)
        
        # 直接用枚举转字典即可,比手动循环构造速度更快
        return dict(enumerate(d))
    
    def _fetch_single_sentiment(self, item):
        """单条数据情感请求逻辑,抽成独立函数供线程池调用"""
        lang = item['lang']
        if lang not in SUPPORTED_LANG:
            return item
        try:
            # 加超时参数避免单条请求卡死
            r = SESSION.post(
                f"{self.API}/sentiment",
                data={'q': item['text'], 'language': lang, 'key': '--'},
                timeout=10
            )
            item['sentiment'] = r.json()
        except Exception as e:
            # 单条请求失败不中断整体流程,标记为空后续补跑即可
            item['sentiment'] = None
            print(f"条目{item['id']}请求失败: {str(e)}")
        return item

    def add_SENTIMENT(self, max_workers=20):
        """
        并发请求情感接口
        :param max_workers: 并发数,根据API限流规则调整,一般10-30即可
        """
        with ThreadPoolExecutor(max_workers=max_workers) as executor:
            # 提交所有任务到线程池
            futures = [
                executor.submit(self._fetch_single_sentiment, value)
                for value in self.json_file.values()
            ]
            # 等待所有任务执行完成
            for future in as_completed(futures):
                future.result()
        return self.json_file

注意:字典是可变对象,线程里对item的修改会直接作用在原始self.json_file的value上,不需要额外做赋值操作。

2. 细节提效点

  • 给请求加异常捕获和重试逻辑:遇到网络波动、API临时限流时自动重试2-3次,避免单条失败导致全量重跑
  • 并发数不要设置过高,否则会触发API限流,建议先从10开始压测,找到不触发限流的最大值
  • 如果数据量超过百万级,不要一次性把全量JSON加载到内存,改成逐块读取、分批处理,处理完一批就写入结果文件,避免内存溢出。

内容的提问来源于stack exchange,提问作者corvusMidnight

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.03 03:33:31