利用Telethon API与Peewee实现Telegram频道新增参与者的获取与增量更新
解决Telegram频道新增参与者的增量存储问题
我来帮你梳理下这个问题,你已经找对了方向,但当前的实现有几个潜在的逻辑漏洞,咱们一步步优化,让增量更新的逻辑更可靠。
先说说你现有代码的问题
你之前用MAX(cudb.channel_usersID)(自增主键最大值)和频道用户数对比的思路是有问题的:
- 自增主键和频道实际用户数没有直接关联,比如用户退出频道、数据库删除记录后,主键最大值和已存储用户数会脱节
- 频道总人数和已存储用户数的差值,不等于新增用户数(可能有用户退出,人数减少但有新用户加入)
- 依赖偏移量计算新增用户的方式,会因为Telegram参与者列表的排序变化(比如用户活跃程度影响排序)导致漏抓或重复抓
更可靠的优化方案:基于用户ID的增量判断
核心思路是:先从数据库获取所有已存储的Telegram用户ID,拉取频道参与者后,直接过滤出不在数据库里的新用户,再进行存储。这种方式不依赖排序或数量对比,完全基于唯一的用户ID判断,是最稳妥的方案。
优化后的完整代码
from telethon import TelegramClient from telethon.tl.functions.channels import GetParticipantsRequest from telethon.tl.types import ChannelParticipantsSearch, ChannelParticipantsRecent from datetime import datetime from schema import channel_users as cudb from dotenv import load_dotenv import os load_dotenv() api_id = os.getenv('API_ID') api_hash = os.getenv('API_HASH') client = TelegramClient('anon', api_id, api_hash) async def get_existing_user_ids(): """从数据库获取所有已存储的Telegram用户ID,存入集合实现O(1)查询""" return set(user.id for user in cudb.select(cudb.id)) async def fetch_all_new_participants(channel): """全量拉取频道参与者,过滤出新增用户(适合中小频道)""" existing_ids = await get_existing_user_ids() new_users = [] offset = 0 limit = 100 while True: participants = await client(GetParticipantsRequest( channel, ChannelParticipantsSearch(''), offset, limit, hash=0 )) if not participants.users: break # 过滤出不在数据库中的新用户 batch_new = [p for p in participants.users if p.id not in existing_ids] new_users.extend(batch_new) offset += len(participants.users) return new_users async def fetch_recent_new_participants(channel, limit=200): """仅拉取最近加入的参与者(适合大频道,减少拉取数据量)""" existing_ids = await get_existing_user_ids() participants = await client(GetParticipantsRequest( channel, ChannelParticipantsRecent(), 0, limit, hash=0 )) return [p for p in participants.users if p.id not in existing_ids] async def save_new_users(users): """批量保存新增用户到数据库""" if not users: print("没有新增用户需要保存") return now = datetime.now().strftime("%d/%m/%Y, %H:%M:%S") for user in users: _, created = cudb.get_or_create( id=user.id, defaults={ 'first_name': user.first_name, 'last_name': user.last_name, 'username': user.username, 'phone': user.phone, 'is_bot': user.bot, 'date_added': now } ) if created: print(f"✅ 新增用户:Telegram ID={user.id},用户名={user.username or '无'}") else: print(f"⚠️ 用户已存在:Telegram ID={user.id}") async def main(): await client.start() my_channel = 'https://t.me/some_channel_url' # 中小频道用全量拉取,大频道建议用fetch_recent_new_participants print("正在获取新增参与者...") new_participants = await fetch_all_new_participants(my_channel) # new_participants = await fetch_recent_new_participants(my_channel, limit=300) await save_new_users(new_participants) with client: client.loop.run_until_complete(main())
方案说明
两种拉取模式
- 全量拉取:适合用户数不多的频道,确保不会遗漏任何新增用户
- 最近用户拉取:用
ChannelParticipantsRecent获取最近加入的一批用户,适合大频道,减少API调用和数据处理量,建议定时(比如每小时)执行一次
核心优化点
- 用集合存储已存在的用户ID,判断新增用户的时间复杂度是O(1),效率极高
- 完全依赖Telegram用户ID作为唯一标识,避免排序、数量变化带来的错误
- 拆分功能为独立函数,代码更清晰,后续维护更方便
额外建议
- 给数据库的
id字段添加唯一索引,确保不会重复存储同一个用户(get_or_create会依赖这个约束) - 如果需要定时执行增量更新,可以结合
schedule库实现定时任务 - 对于超大型频道,可以考虑分页拉取时记录上次拉取到的最大用户ID,下次从该ID开始过滤,但这种方式需要确保列表是按ID排序的,不如直接用最近用户拉取稳妥
内容的提问来源于stack exchange,提问作者John
相关产品推荐
相关产品推荐

