Telethon爬取Telegram数据报错排查:请求失败与会话ID错误
Telegram爬取服务器响应慢及报错问题排查
问题描述
使用Telethon进行Telegram数据爬取时,多数情况服务器响应过慢,频繁报错,但相同代码偶尔能正常返回数据。不清楚问题根源,以下是代码及遇到的错误:
代码示例
from telethon.sync import TelegramClient import datetime import pandas as pd import pymongo api_id = xxxxxxx api_hash = 'mycorrect_api_hash' chats = ['group-of-telegram-here'] clientd = pymongo.MongoClient("mongodb://localhost:27017") db = clientd['xxxx'] collection = db['mycollection'] my_list = [] for chat in chats: with TelegramClient('mysession', api_id, api_hash) as client: for message in client.iter_messages(chat, offset_date=datetime.date(2023, 1, 11), reverse=True): print(message) my_list.append({"group": chat, "sender": message.sender_id, "text": message.text, "date": message.date}) collection.insert_many(my_list)
常见错误
Request was unsuccessful 6 time(s)
Security error while unpacking a received message: Server replied with a wrong session ID
问题分析与优化方案
错误原因解析
Request was unsuccessful:请求重试多次失败,大概率是请求频率过高触发Telegram限流,或是网络波动导致连接不稳定。Wrong session ID:循环内反复创建TelegramClient实例,会重复初始化会话,极易引发会话冲突,导致服务器返回会话ID不匹配的安全错误。
代码优化点
- 复用客户端实例:不要在
for chat循环内重复创建客户端,一次性初始化即可,避免会话冲突。 - 处理限流异常:Telethon内置
FloodWaitError,捕获后按要求等待即可自动恢复请求。 - 分批插入数据库:大数量数据一次性插入易导致内存溢出,分批插入更稳定。
- 添加合理延迟:降低请求频率,减少触发限流的概率。
优化后代码
from telethon.sync import TelegramClient from telethon.errors import FloodWaitError import datetime import pymongo import time api_id = xxxxxxx api_hash = 'mycorrect_api_hash' chats = ['group-of-telegram-here'] clientd = pymongo.MongoClient("mongodb://localhost:27017") db = clientd['xxxx'] collection = db['mycollection'] my_list = [] batch_size = 100 # 每100条数据插入一次 # 只初始化一次客户端 with TelegramClient('mysession', api_id, api_hash) as client: for chat in chats: try: for message in client.iter_messages(chat, offset_date=datetime.date(2023, 1, 11), reverse=True): print(message) my_list.append({ "group": chat, "sender": message.sender_id, "text": message.text, "date": message.date }) # 达到批次大小就插入数据库 if len(my_list) >= batch_size: collection.insert_many(my_list) my_list.clear() time.sleep(1) # 插入后短暂延迟 # 插入剩余未达批次的数据 if my_list: collection.insert_many(my_list) except FloodWaitError as e: # 遇到限流,按Telegram要求等待后重试 print(f"触发限流,等待{e.seconds}秒") time.sleep(e.seconds) continue except Exception as e: print(f"发生错误: {str(e)}") time.sleep(5) # 其他错误等待5秒后继续
额外建议
- 确认
api_id和api_hash正确,且爬取用的Telegram账号无限制记录。 - 若服务器在国内,需确保网络能稳定访问Telegram,必要时配置代理(Telethon支持
socks5等代理类型)。 - 避免短时间内爬取超大量数据,严格遵循Telegram的API使用规范,防止账号被封禁。
内容的提问来源于stack exchange,提问作者Ziviz Zaz
相关产品推荐
相关产品推荐

