Python读取双文本文件作为API参数调用的高效实现问题
高效读取双文件并调用API的优化方案
嘿,看你的代码和需求,你是想读取两个文本文件里的地址和城市数据,处理后传给API调用,但在高效实现上遇到了卡点对吧?我来帮你梳理下现有代码的问题,再给出更靠谱的优化方案~
先聊聊现有代码的几个小问题
- 没有用
with语句管理文件句柄:直接open文件后如果程序中途出错,容易导致文件资源泄漏,而且手动关闭也容易忘。 - 分开遍历两个列表会覆盖变量:你现在的两个
for循环分别处理地址和城市,最后self.k和self.j只会保留最后一行的处理结果,没法把每一组地址和城市对应起来传给API。 - 数据预处理可以更紧凑:重复的分割、清理逻辑可以复用,减少冗余代码。
优化方案1:安全高效的文件读取+同步API调用
先解决文件读取的资源管理问题,再把地址和城市一一配对处理,最后调用API:
import re import requests class YourAPIClient: def __init__(self): self.addresses = [] self.cities = [] def read_addresses(self): # 用with语句自动管理文件,无需手动close,安全又简洁 with open("addresses.txt", "r") as f, open("cities.txt", "r") as f2: # 读取时直接去掉换行符,避免后续处理麻烦 self.addresses = [line.strip() for line in f] self.cities = [line.strip() for line in f2] def _clean_text(self, text): # 把重复的清理逻辑抽成私有方法,复用性更强 sep = ' Placeholder' cleaned = text.split(sep, 1)[0] cleaned = re.sub('\s+',' ', cleaned).strip() return cleaned def get_data(self): # 先检查两个文件的行数是否一致,避免配对错误 if len(self.addresses) != len(self.cities): raise ValueError("地址文件和城市文件的行数不一致,请检查!") # 一一配对处理数据 for addr, city in zip(self.addresses, self.cities): cleaned_addr = self._clean_text(addr) cleaned_city = self._clean_text(city) # 替换成你的API调用逻辑,这里用requests举例 response = requests.get( "https://your-api-endpoint.com", params={"address": cleaned_addr, "city": cleaned_city} ) # 处理API响应,比如保存结果、打印日志等 self._handle_response(response) def _handle_response(self, response): if response.status_code == 200: print("API请求成功:", response.json()) # 这里可以添加保存结果到数据库/文件的逻辑 else: print(f"API请求失败,状态码: {response.status_code}")
优化方案2:大文件场景下的逐行处理
如果你的文本文件特别大,一次性把所有内容读到内存里会占用过多资源,这时候可以逐行读取处理,不用把所有数据存在内存中:
def read_and_process_large_files(self): with open("addresses.txt", "r") as f, open("cities.txt", "r") as f2: # 逐行配对读取,内存占用极低 for addr_line, city_line in zip(f, f2): addr = addr_line.strip() city = city_line.strip() cleaned_addr = self._clean_text(addr) cleaned_city = self._clean_text(city) # 直接调用API,不用缓存所有数据 self._call_api(cleaned_addr, cleaned_city)
优化方案3:异步API调用(提升IO密集型场景效率)
如果API调用是耗时的IO操作,用异步请求可以大幅提升效率,避免等待单个请求完成再发下一个:
import aiohttp import asyncio class AsyncAPIClient(YourAPIClient): async def _fetch_api(self, session, addr, city): async with session.get( "https://your-api-endpoint.com", params={"address": addr, "city": cleaned_city} ) as response: return await response.json() async def get_data_async(self): if len(self.addresses) != len(self.cities): raise ValueError("地址文件和城市文件的行数不一致,请检查!") processed_pairs = [] for addr, city in zip(self.addresses, self.cities): processed_pairs.append((self._clean_text(addr), self._clean_text(city))) # 批量异步请求 async with aiohttp.ClientSession() as session: tasks = [self._fetch_api(session, addr, city) for addr, city in processed_pairs] responses = await asyncio.gather(*tasks) # 批量处理响应 for resp in responses: print("异步API响应:", resp)
额外注意点
- 一定要添加异常处理:比如文件不存在、API请求超时、网络错误等,让程序更健壮。
- 如果API支持批量请求,尽量把多组数据打包成一个请求发送,比逐个调用效率更高。
内容的提问来源于stack exchange,提问作者rustyshackleford
相关产品推荐
相关产品推荐

