You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提升Python脚本速度?处理HTTP请求修饰符与JSON追加写入

Python脚本优化思路:保留进度条同时提速4500条数据处理

Hey there! Let's break down how to speed up your script while keeping that handy progress bar you find useful. With 4500 entries to process, small tweaks can make a huge difference in runtime. Here are practical, actionable optimizations:

1. 缩小Try-Except块范围,减少异常处理开销

Try-except blocks do have a performance cost when executed frequently—especially if your else clause runs every time (since exceptions are likely rare in your case). To cut this down:

  • Wrap only error-prone code: Don't stuff your entire processing logic into the try block. If only the HTTP request can fail, isolate that part. For example:
    for id in tqdm(idents):
        # 非风险逻辑放在try块外
        modifier = f"{id}@example.com"
        request_payload = build_request_data(modifier)
        
        # 仅将HTTP请求包在try-except中
        try:
            response = requests.post(target_url, data=request_payload)
            response.raise_for_status()  # 主动触发HTTP错误异常
        except requests.exceptions.RequestException as e:
            print(f"Failed to process {id}: {e}")
            continue
        else:
            # 处理成功响应的逻辑
            result = parse_response(response)
            results.append(result)
    
  • 前置校验减少异常触发: Validate idents upfront to skip invalid entries before they even reach the try block. For example, if your modifier requires exactly 5 characters:
    valid_idents = [id for id in idents if len(id) == 5 and id.isalnum()]
    for id in tqdm(valid_idents):
        # 处理逻辑
    

2. 批量写入JSON,避免频繁磁盘IO

Appending to a JSON file after every single entry is a major slowdown—disk IO is one of the slowest operations in computing. Instead:

  • Collect results in memory first: Store processed results in a list, then write everything to the file once at the end. If you're worried about data loss if the script crashes, split into batches (e.g., write every 500 entries):
    results = []
    batch_size = 500
    
    for idx, id in enumerate(tqdm(idents)):
        # 处理id并生成result
        results.append(result)
        
        # 每batch_size条写入一次
        if (idx + 1) % batch_size == 0:
            with open("output.json", "a") as f:
                json.dump(results, f)
                f.write("\n")  # 可选:让每条结果单独一行,方便后续读取
            results = []
    
    # 写入剩余的结果
    if results:
        with open("output.json", "a") as f:
            json.dump(results, f)
    

This cuts 4500 IO operations down to just 9 (for a 500-entry batch)—a massive speed boost.

3. 优化HTTP请求效率

Chances are, the HTTP requests themselves are the biggest bottleneck, not the try-except blocks. Here's how to speed them up:

  • Use a session for connection pooling: requests.Session() reuses TCP connections between requests, eliminating the overhead of establishing a new connection every time:
    session = requests.Session()
    for id in tqdm(idents):
        # 使用session发送请求,复用连接
        response = session.post(target_url, data=request_payload)
    
  • Parallelize requests: Sequential requests waste tons of time waiting for responses. Use thread pooling (simple to implement) to process multiple requests at once:
    from concurrent.futures import ThreadPoolExecutor
    from tqdm import tqdm
    
    session = requests.Session()
    max_workers = 10  # 根据目标网站承受能力调整,避免被封禁
    
    def process_single_id(id):
        # 单个id的完整处理逻辑
        modifier = f"{id}@example.com"
        try:
            response = session.post(target_url, data=build_request_data(modifier))
            response.raise_for_status()
            return parse_response(response)
        except Exception as e:
            print(f"Error with {id}: {e}")
            return None
    
    with ThreadPoolExecutor(max_workers=max_workers) as executor:
        # 用tqdm显示并行处理的进度
        results = list(tqdm(executor.map(process_single_id, idents), total=len(idents)))
    
    # 过滤无效结果并写入文件
    valid_results = [res for res in results if res is not None]
    with open("output.json", "w") as f:
        json.dump(valid_results, f)
    

4. 替换为轻量进度条

If your custom Progressbar() is heavy on resources, switch to tqdm—it's designed to have minimal overhead while providing clear progress updates. It integrates seamlessly with both sequential and parallel loops, as shown in the examples above.

5. 消除循环内的冗余操作

Check for any code that runs inside the loop but doesn't need to:

  • Move static objects (like request headers, target URL definitions) outside the loop—don't redefine them every iteration.
  • Avoid repeated calculations (e.g., if you're formatting the modifier the same way every time, ensure you're not doing unnecessary string operations).

内容的提问来源于stack exchange,提问作者Robert Millard

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 10:34:17