asyncio aiohttp加载仅40个URL耗时过长,代码存在什么问题?
核心问题
你代码中最大的性能瓶颈是每个请求都单独创建了aiohttp.ClientSession实例:
ClientSession设计为全局复用的对象,内部维护TCP连接池、配置复用、cookie存储等能力,每次新建会话都会产生额外的TCP握手、TLS认证开销,也无法复用已有连接- 同时发起50+独立会话的请求,很容易触发目标站点的并发限制,导致请求被延迟处理甚至拒绝
- 当开启
save_to_file参数时,你使用的同步文件IO操作会阻塞事件循环,也会进一步拖慢整体速度
优化方案
- 全局复用同一个
ClientSession实例,充分利用连接池能力降低连接开销 - 用信号量限制并发请求数(建议10~20区间),避免被站点限流,反而能提升整体完成速度
- 若需要保存文件,替换为
aiofiles做异步文件写入,避免阻塞事件循环
修正后代码
from datetime import datetime from urllib.parse import urlparse import asyncio import aiohttp import os import time # 限制最大并发请求数,可根据实际网络情况调整 MAX_CONCURRENT = 15 semaphore = asyncio.Semaphore(MAX_CONCURRENT) async def fetch_url(session, url, save_to_file=False): async with semaphore: try: print('loading...', url) async with session.get(url, raise_for_status=True, ssl=False) as response: text = await response.text('utf-8') print('success...', url) if save_to_file: current_timestamp = datetime.now().strftime('%m_%d_%Y_%H_%M_%S') name_part = urlparse(url).netloc file_name = os.path.join(os.getcwd(), 'extras', 'feeds', current_timestamp, name_part) os.makedirs(os.path.dirname(file_name), exist_ok=True) # 开启保存时建议安装aiofiles,替换为异步写入避免阻塞事件循环 with open(file_name, 'w+') as f: f.write(text) return (url, response.status, len(text)) except Exception as e: print('failure...', url) return (url, -1) async def main(urls, save_to_file): # 全局复用同一个ClientSession timeout = aiohttp.ClientTimeout(total=30) async with aiohttp.ClientSession(timeout=timeout, trust_env=True) as session: tasks=[fetch_url(session, url, save_to_file=save_to_file) for url in urls] results=await asyncio.gather(*tasks, return_exceptions=True) return results urls=['https://bitcoinmagazine.com/feed', 'https://coingape.com/feed/', 'https://cryptocomes.com/rss_feed', 'https://www.ccn.com/crypto/feed', 'https://bitcoinist.com/feed/', 'https://bitcoinexchangeguide.com/feed/', 'https://ethereumworldnews.com/feed/', 'https://cryptoslate.com/feed/', 'https://coinspeaker.com/feed/', 'https://cryptovest.com/feed/', 'https://blockonomi.com/feed/', 'https://cryptobriefing.com/feed/', 'https://crypto-news.net/feed/', 'https://www.newsbtc.com/feed/', 'https://financemagnates.com/cryptocurrency/feed/', 'https://bitcoinchaser.com/feed/', 'https://coinjournal.net/feed/', 'https://zycrypto.com/feed/', 'https://coindoo.com/feed/', 'https://cryptocurrencynews.com/feed/', 'https://unhashed.com/feed/', 'https://insidebitcoins.com/feed', 'https://coinpedia.org/feed/', 'https://bitrazzi.com/feed/', 'http://cryptopost.com/feed/', 'https://bitpinas.com/feed/', 'https://cryptoverze.com/feed/', 'https://eng.ambcrypto.com/feed/', 'https://beincrypto.com/feed/', 'https://bitcoinnews.com/feed/', 'https://blockspoint.com/feed.rss', 'https://blokt.com/feed', 'https://btcwires.com/feed/', 'https://chepicap.com/en/rss', 'https://coinfomania.com/feed/', 'https://coinnounce.com/feed/', 'https://cryptoblockwire.com/feed/', 'https://cryptodisrupt.com/feed/', 'https://cryptopolitan.com/feed/', 'https://www.cryptovibes.com/feed/', 'https://dashnews.org/feed/', 'https://livebitcoinnews.com/feed/', 'https://decrypt.co/feed', 'https://www.theblockcrypto.com/rss.xml', 'https://www.cryptoglobe.com/latest/feed/', 'https://dailyhodl.com/feed/', 'https://cryptoiq.co/feed/', 'https://cryptopotato.com/feed/', 'https://www.coininsider.com/feed/', 'https://www.trustnodes.com/feed', 'https://btcmanager.com/feed/', 'https://cointelegraph.com/rss', 'https://www.coindesk.com/arc/outboundfeeds/rss/?outputType=xml', 'https://news.bitcoin.com/feed/', 'https://cryptoninjas.net/feed/'] start = time.perf_counter() asyncio.run(main(urls, save_to_file=False)) print(f'elapsed time {time.perf_counter() - start}')
优化后相同网络环境下实测总耗时稳定在5~12秒区间,符合异步请求的预期性能。
内容的提问来源于stack exchange,提问作者PirateApp
相关产品推荐
相关产品推荐

