You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python读取双文本文件作为API参数调用的高效实现问题

高效读取双文件并调用API的优化方案

嘿,看你的代码和需求,你是想读取两个文本文件里的地址和城市数据,处理后传给API调用,但在高效实现上遇到了卡点对吧?我来帮你梳理下现有代码的问题,再给出更靠谱的优化方案~

先聊聊现有代码的几个小问题

  • 没有用with语句管理文件句柄:直接open文件后如果程序中途出错,容易导致文件资源泄漏,而且手动关闭也容易忘。
  • 分开遍历两个列表会覆盖变量:你现在的两个for循环分别处理地址和城市,最后self.k和self.j只会保留最后一行的处理结果,没法把每一组地址和城市对应起来传给API。
  • 数据预处理可以更紧凑:重复的分割、清理逻辑可以复用,减少冗余代码。

优化方案1:安全高效的文件读取+同步API调用

先解决文件读取的资源管理问题,再把地址和城市一一配对处理,最后调用API:

import re
import requests

class YourAPIClient:
    def __init__(self):
        self.addresses = []
        self.cities = []

    def read_addresses(self):
        # 用with语句自动管理文件,无需手动close,安全又简洁
        with open("addresses.txt", "r") as f, open("cities.txt", "r") as f2:
            # 读取时直接去掉换行符,避免后续处理麻烦
            self.addresses = [line.strip() for line in f]
            self.cities = [line.strip() for line in f2]

    def _clean_text(self, text):
        # 把重复的清理逻辑抽成私有方法,复用性更强
        sep = ' Placeholder'
        cleaned = text.split(sep, 1)[0]
        cleaned = re.sub('\s+',' ', cleaned).strip()
        return cleaned

    def get_data(self):
        # 先检查两个文件的行数是否一致,避免配对错误
        if len(self.addresses) != len(self.cities):
            raise ValueError("地址文件和城市文件的行数不一致,请检查!")
        
        # 一一配对处理数据
        for addr, city in zip(self.addresses, self.cities):
            cleaned_addr = self._clean_text(addr)
            cleaned_city = self._clean_text(city)
            
            # 替换成你的API调用逻辑,这里用requests举例
            response = requests.get(
                "https://your-api-endpoint.com",
                params={"address": cleaned_addr, "city": cleaned_city}
            )
            # 处理API响应,比如保存结果、打印日志等
            self._handle_response(response)

    def _handle_response(self, response):
        if response.status_code == 200:
            print("API请求成功:", response.json())
            # 这里可以添加保存结果到数据库/文件的逻辑
        else:
            print(f"API请求失败,状态码: {response.status_code}")

优化方案2:大文件场景下的逐行处理

如果你的文本文件特别大,一次性把所有内容读到内存里会占用过多资源,这时候可以逐行读取处理,不用把所有数据存在内存中:

def read_and_process_large_files(self):
    with open("addresses.txt", "r") as f, open("cities.txt", "r") as f2:
        # 逐行配对读取,内存占用极低
        for addr_line, city_line in zip(f, f2):
            addr = addr_line.strip()
            city = city_line.strip()
            
            cleaned_addr = self._clean_text(addr)
            cleaned_city = self._clean_text(city)
            
            # 直接调用API,不用缓存所有数据
            self._call_api(cleaned_addr, cleaned_city)

优化方案3:异步API调用(提升IO密集型场景效率)

如果API调用是耗时的IO操作,用异步请求可以大幅提升效率,避免等待单个请求完成再发下一个:

import aiohttp
import asyncio

class AsyncAPIClient(YourAPIClient):
    async def _fetch_api(self, session, addr, city):
        async with session.get(
            "https://your-api-endpoint.com",
            params={"address": addr, "city": cleaned_city}
        ) as response:
            return await response.json()

    async def get_data_async(self):
        if len(self.addresses) != len(self.cities):
            raise ValueError("地址文件和城市文件的行数不一致,请检查!")
        
        processed_pairs = []
        for addr, city in zip(self.addresses, self.cities):
            processed_pairs.append((self._clean_text(addr), self._clean_text(city)))
        
        # 批量异步请求
        async with aiohttp.ClientSession() as session:
            tasks = [self._fetch_api(session, addr, city) for addr, city in processed_pairs]
            responses = await asyncio.gather(*tasks)
            
            # 批量处理响应
            for resp in responses:
                print("异步API响应:", resp)

额外注意点

  • 一定要添加异常处理:比如文件不存在、API请求超时、网络错误等,让程序更健壮。
  • 如果API支持批量请求,尽量把多组数据打包成一个请求发送,比逐个调用效率更高。

内容的提问来源于stack exchange,提问作者rustyshackleford

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 12:07:44