You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Telethon爬取Telegram数据报错排查:请求失败与会话ID错误

Telegram爬取服务器响应慢及报错问题排查

问题描述

使用Telethon进行Telegram数据爬取时,多数情况服务器响应过慢,频繁报错,但相同代码偶尔能正常返回数据。不清楚问题根源,以下是代码及遇到的错误:

代码示例

from telethon.sync import TelegramClient
import datetime
import pandas as pd
import pymongo
api_id = xxxxxxx
api_hash = 'mycorrect_api_hash'
chats = ['group-of-telegram-here']
clientd = pymongo.MongoClient("mongodb://localhost:27017")
db = clientd['xxxx']
collection = db['mycollection']
my_list = []
for chat in chats:
    with TelegramClient('mysession', api_id, api_hash) as client:
        for message in client.iter_messages(chat, offset_date=datetime.date(2023, 1, 11), reverse=True):
            print(message)
            my_list.append({"group": chat, "sender": message.sender_id, "text": message.text, "date": message.date})

collection.insert_many(my_list)

常见错误

Request was unsuccessful 6 time(s)
Security error while unpacking a received message: Server replied with a wrong session ID

问题分析与优化方案

错误原因解析

  • Request was unsuccessful:请求重试多次失败,大概率是请求频率过高触发Telegram限流,或是网络波动导致连接不稳定。
  • Wrong session ID:循环内反复创建TelegramClient实例,会重复初始化会话,极易引发会话冲突,导致服务器返回会话ID不匹配的安全错误。

代码优化点

  1. 复用客户端实例:不要在for chat循环内重复创建客户端,一次性初始化即可,避免会话冲突。
  2. 处理限流异常:Telethon内置FloodWaitError,捕获后按要求等待即可自动恢复请求。
  3. 分批插入数据库:大数量数据一次性插入易导致内存溢出,分批插入更稳定。
  4. 添加合理延迟:降低请求频率,减少触发限流的概率。

优化后代码

from telethon.sync import TelegramClient
from telethon.errors import FloodWaitError
import datetime
import pymongo
import time

api_id = xxxxxxx
api_hash = 'mycorrect_api_hash'
chats = ['group-of-telegram-here']
clientd = pymongo.MongoClient("mongodb://localhost:27017")
db = clientd['xxxx']
collection = db['mycollection']
my_list = []
batch_size = 100  # 每100条数据插入一次

# 只初始化一次客户端
with TelegramClient('mysession', api_id, api_hash) as client:
    for chat in chats:
        try:
            for message in client.iter_messages(chat, offset_date=datetime.date(2023, 1, 11), reverse=True):
                print(message)
                my_list.append({
                    "group": chat,
                    "sender": message.sender_id,
                    "text": message.text,
                    "date": message.date
                })
                # 达到批次大小就插入数据库
                if len(my_list) >= batch_size:
                    collection.insert_many(my_list)
                    my_list.clear()
                    time.sleep(1)  # 插入后短暂延迟
            
            # 插入剩余未达批次的数据
            if my_list:
                collection.insert_many(my_list)
        
        except FloodWaitError as e:
            # 遇到限流,按Telegram要求等待后重试
            print(f"触发限流,等待{e.seconds}秒")
            time.sleep(e.seconds)
            continue
        
        except Exception as e:
            print(f"发生错误: {str(e)}")
            time.sleep(5)  # 其他错误等待5秒后继续

额外建议

  • 确认api_id和api_hash正确,且爬取用的Telegram账号无限制记录。
  • 若服务器在国内,需确保网络能稳定访问Telegram,必要时配置代理(Telethon支持socks5等代理类型)。
  • 避免短时间内爬取超大量数据,严格遵循Telegram的API使用规范,防止账号被封禁。

内容的提问来源于stack exchange,提问作者Ziviz Zaz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 11:01:34