通过WhatsApp Cloud API发10万条消息时TCP连接挂起问题排查
疑问解答
Webhook未及时响应不会直接引发发送服务器的TCP连接挂起。发送消息是你的服务器向Meta的WhatsApp API发起请求,而Webhook是Meta向你的通知服务器推送事件,这两个是完全独立的通信链路,彼此的状态不会直接影响。你的发送端连接挂起问题,根源在发送逻辑本身。
可行解决方案
一、优化发送逻辑的并发与资源控制
避免一次性创建全量任务
不要把10万个号码一次性塞入异步任务队列,改为分批次处理(比如每批次50-100条),批次之间添加短暂延迟,防止瞬间压爆连接池或触发Meta的限流机制。精确控制并发请求数
用asyncio.Semaphore替代单纯依赖TCP连接池限制,确保同时运行的请求数在合理范围,避免连接资源耗尽。同时调整TCP连接池大小略大于并发数,预留缓冲空间。添加请求超时机制
在请求Meta API时设置超时时间,防止单个请求长时间卡住占用连接,导致后续请求无法获取连接而挂起。
修改后的示例代码:
import asyncio import aiohttp import ujson async def send_msg(phone: str, session: aiohttp.ClientSession, semaphore: asyncio.Semaphore): async with semaphore: json_payload = { "messaging_product": "whatsapp", "recipient_type": "individual", "to": phone, "type": 'template', "template": { "name": 'template_123', "language": {"code": 'ar'}, "components": [] } } headers = {'Authorization': f'Bearer {whatsapp_api_key}'} try: async with session.post( f'https://graph.facebook.com/v14.0/{campaign.user.phone_number_id}/messages', json=json_payload, headers=headers, timeout=aiohttp.ClientTimeout(total=10) # 设置10秒超时 ) as response: print(f"{phone} - 状态码: {response.status}") # 处理限流情况 if response.status == 429: retry_after = int(response.headers.get('Retry-After', 5)) await asyncio.sleep(retry_after) await send_msg(phone, session, semaphore) except Exception as e: print(f"{phone} 发送失败: {str(e)}") # 可选:添加重试逻辑,比如最多重试3次 # await asyncio.sleep(2) # await send_msg(phone, session, semaphore) async def send_all(phones: list): semaphore = asyncio.Semaphore(20) # 限制同时20个请求 my_conn = aiohttp.TCPConnector(limit=30) # 连接池大小略大于并发数 async with aiohttp.ClientSession(connector=my_conn, json_serialize=ujson.dumps) as session: batch_size = 50 for i in range(0, len(phones), batch_size): batch = phones[i:i+batch_size] tasks = [asyncio.ensure_future(send_msg(phone, session, semaphore)) for phone in batch] await asyncio.gather(*tasks, return_exceptions=True) await asyncio.sleep(1) # 批次间延迟1秒 start = time.time() asyncio.run(send_all(data)) end = time.time()
二、严格处理Meta API的限流规则
Meta的WhatsApp API有明确的速率限制,频繁触发429限流会导致连接被临时封禁,进而引发TCP挂起:
- 必须捕获429状态码,根据返回的
Retry-After头延迟重试 - 主动控制发送速率,比如每秒发送20-30条(不同账号的限流阈值不同,可参考Meta官方文档调整)
三、修复Webhook服务器的响应问题
虽然Webhook不直接影响发送,但长期无法及时返回200会影响账号信誉,甚至导致Meta停止推送通知:
- Webhook收到请求后优先返回200状态码,再异步处理通知内容,不要在响应链路中做耗时操作
- 确保Webhook在10秒内完成响应(Meta要求20秒内,留足冗余)
- 对Webhook服务器进行资源扩容或添加负载均衡,分散请求压力
四、排查TCP连接挂起的具体原因
- 查看发送服务器的系统日志(如
dmesg、/var/log/syslog),排查TCP连接相关错误 - 使用
tcpdump抓包,分析与Meta API的连接状态,判断是本地连接耗尽还是Meta主动断开 - 开启aiohttp的debug日志,查看连接池的使用情况:
import logging logging.basicConfig(level=logging.DEBUG)
内容的提问来源于stack exchange,提问作者Ahmed Zeini
相关产品推荐
相关产品推荐

