Python3.7通过API批量取数遇502错误,求优化方案
高效获取UMLS CUI数据的优化方案
问题根源
502错误大概率是高频单线程请求触发API限流,或是长时间单线程请求导致的连接超时。UMLS API存在明确的请求速率限制,12000次请求后触发限流是典型的超限表现。
优化方案
1. 增加重试机制
针对502、503这类临时错误自动重试,避免流程中断,同时用指数退避策略降低重试冲突概率。
2. 严格限流控速
遵守UMLS API的速率限制(通常为每秒2-5次请求),添加固定延迟或令牌桶算法控制请求频率。
3. 复用连接池
用requests.Session复用TCP连接,减少握手开销,提升请求效率。
4. 异步批量请求(最优选择)
通过异步框架同时处理多个请求,结合批量查询端点一次提交多个CUI,大幅减少请求次数,提升整体效率。
完整优化代码示例
方案一:带重试和限流的单线程优化
import requests import time from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type # 配置参数 API_KEY = "aaaaaaaaaaaaaaaaaaaaaaaa" BASE_URL = "https://uts-ws.nlm.nih.gov/rest/content/current/CUI/" RATE_LIMIT = 2 # 每秒请求数,按UMLS文档调整 RETRY_MAX_ATTEMPTS = 3 # 初始化文件句柄 umls_cui_file = open('umls_cui_names.txt', 'w') missed_cui_file = open('not_found_cui.txt', 'w') # 创建会话复用连接池 session = requests.Session() session.headers.update({'Content-Type': 'application/json'}) session.params.update({'apiKey': API_KEY}) @retry( stop=stop_after_attempt(RETRY_MAX_ATTEMPTS), wait=wait_exponential(multiplier=1, min=2, max=10), retry=retry_if_exception_type((requests.exceptions.HTTPError, requests.exceptions.ConnectionError)) ) def get_cui(cui): url = f"{BASE_URL}{cui}" response = session.get(url) response.raise_for_status() data = response.json() name = data['result']['name'] print(f"{cui}\t{name}") umls_cui_file.write(f"{cui}\t{name}\n") umls_cui_file.flush() # 实时写入避免数据丢失 # 遍历处理所有CUI for idx, cui in enumerate(df_cui['CUI'], 1): print(f"处理进度: {idx}/{len(df_cui)} - {cui}") try: get_cui(cui) except Exception as e: print(f"{cui} 获取失败: {str(e)}") missed_cui_file.write(f"{cui}\n") missed_cui_file.flush() # 限流延迟 time.sleep(1/RATE_LIMIT) # 清理资源 session.close() umls_cui_file.close() missed_cui_file.close()
方案二:异步批量请求(效率翻倍)
import aiohttp import asyncio from tenacity import retry, stop_after_attempt, wait_exponential API_KEY = "aaaaaaaaaaaaaaaaaaaaaaaa" BASE_BATCH_URL = "https://uts-ws.nlm.nih.gov/rest/content/current/CUIs" CONCURRENT_LIMIT = 5 # 并发数,按API限制调整 RETRY_MAX_ATTEMPTS = 3 BATCH_SIZE = 10 # 每批提交的CUI数量,按API上限调整 umls_cui_file = open('umls_cui_names.txt', 'w') missed_cui_file = open('not_found_cui.txt', 'w') async def fetch_batch(session, cui_batch): params = { 'apiKey': API_KEY, 'cui': ','.join(cui_batch) } @retry( stop=stop_after_attempt(RETRY_MAX_ATTEMPTS), wait=wait_exponential(multiplier=1, min=2, max=10) ) async def _fetch(): async with session.get(BASE_BATCH_URL, params=params) as response: response.raise_for_status() data = await response.json() # 处理成功返回的CUI for result in data['result']: cui = result['ui'] name = result['name'] print(f"{cui}\t{name}") umls_cui_file.write(f"{cui}\t{name}\n") umls_cui_file.flush() # 处理未找到的CUI(需根据API返回结构调整) if 'errors' in data: for error in data['errors']: missed_cui = error['ui'] missed_cui_file.write(f"{missed_cui}\n") missed_cui_file.flush() await _fetch() async def main(): cui_list = df_cui['CUI'].tolist() # 拆分CUI为批量 batches = [cui_list[i:i+BATCH_SIZE] for i in range(0, len(cui_list), BATCH_SIZE)] async with aiohttp.ClientSession(headers={'Content-Type': 'application/json'}) as session: # 用信号量控制并发数 semaphore = asyncio.Semaphore(CONCURRENT_LIMIT) async def bounded_fetch(batch): async with semaphore: await fetch_batch(session, batch) # 批量启动任务 tasks = [bounded_fetch(batch) for batch in batches] await asyncio.gather(*tasks) if __name__ == "__main__": asyncio.run(main()) umls_cui_file.close() missed_cui_file.close()
额外注意事项
- 务必查阅UMLS API官方文档,确认速率限制和批量查询的参数格式,避免参数错误。
- 实时调用
flush()写入文件,防止程序崩溃导致数据丢失。 - 单独记录失败的CUI,最后单独处理这些请求。
- 若API要求使用OAuth令牌而非固定API Key,需先获取令牌并定期刷新。
内容的提问来源于stack exchange,提问作者rshar
相关产品推荐
相关产品推荐

