You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修改Python代码实现Discord API批量数据拉取并全量存入DataFrame

问题原因
  • 变量名严重冲突:一是外层循环遍历频道ID的变量i和内层遍历消息的变量i重复,导致值被覆盖;二是循环内部定义了名为channel_id的列表,直接覆盖了最开始存储所有频道ID的全局列表,后续循环无法拿到正确的频道ID
  • 集合转换打乱对应关系:使用zip(set(collection_name), set(channel_id))遍历,set是无序去重结构,会完全打乱名称和频道ID原本的一一对应关系,数据量越大匹配错误率越高
  • 存储逻辑错误:每个频道的结果列表、discord_dict都定义在循环内部,每次循环都会清空上一次的结果,且DataFrame生成、CSV导出逻辑放在for循环外,只会保留最后一次循环的频道数据
  • 无异常处理:Discord API有严格的限流规则,上千组请求时很容易出现请求失败、返回格式异常的问题,没有处理逻辑会直接导致程序中断
  • 冗余的while循环:写了while True但直接加了break,没有实际作用
修改后代码
import requests
import json
import pandas as pd
import time

# 配置项
collection_name = ['A','B','C']
channel_id_list = ['880281356335206440','899531896571187210','293431696871187210']  
headers = {
    # 请补充你的Discord请求头,比如Authorization等
    "Authorization": "你的Discord Token",
    "Content-Type": "application/json"
}
REQUEST_RETRY_COUNT = 3 # 单次请求失败重试次数
REQUEST_INTERVAL = 1 # 请求间隔(秒),适配Discord限流规则,可根据实际情况调整
SAVE_SEPARATE_CSV = False # 是否每个频道单独存CSV,False则所有数据合并存总表

# 存储所有数据的总列表
all_records = []

# 直接遍历原始列表,保证名称和ID一一对应
for c, channel_id in zip(collection_name, channel_id_list):
    print(f"正在处理频道:{c}(ID:{channel_id})")
    # 请求重试逻辑
    retry_num = 0
    json_texts = None
    while retry_num < REQUEST_RETRY_COUNT:
        try:
            r = requests.get(f'https://discord.com/api/v9/guilds/{channel_id}/messages/search?content=buy', headers=headers, timeout=10)
            if r.status_code == 200:
                json_texts = json.loads(r.text)
                break
            elif r.status_code == 429:
                # 触发限流,等待响应头指定的时间后重试
                retry_after = int(r.headers.get('Retry-After', 5))
                print(f"触发限流,等待{retry_after}秒后重试")
                time.sleep(retry_after)
                retry_num +=1
            else:
                print(f"请求失败,状态码:{r.status_code},第{retry_num+1}次重试")
                time.sleep(REQUEST_INTERVAL)
                retry_num +=1
        except Exception as e:
            print(f"请求异常:{str(e)},第{retry_num+1}次重试")
            time.sleep(REQUEST_INTERVAL)
            retry_num +=1
    if not json_texts or 'messages' not in json_texts:
        print(f"频道{c}无有效返回数据,跳过")
        continue
    # 遍历当前频道的消息
    current_channel_records = []
    for msg_arr in json_texts['messages']:
        msg = msg_arr[0]
        record = {
            'id': msg['id'],
            'type': msg['type'],
            'content': msg['content'],
            'channel_id': msg['channel_id'],
            'attachments': msg['attachments'],
            'embeds': msg['embeds'],
            'mention_roles': msg['mention_roles'],
            'pinned': msg['pinned'],
            'mention_everyone': msg['mention_everyone'],
            'tts': msg['tts'],
            'timestamp': msg['timestamp'],
            'edited_timestamp': msg['edited_timestamp'],
            'flags': msg['flags'],
            'components': msg['components'],
            'author': msg['author'],
            'mentions': msg['mentions'],
            'collectionName': c
        }
        current_channel_records.append(record)
        time.sleep(0.01)
    # 存储逻辑
    if SAVE_SEPARATE_CSV:
        # 每个频道单独存CSV
        df = pd.DataFrame(current_channel_records)
        df.to_csv(f'{c}_buy.csv', index=False, encoding='utf-8-sig')
    else:
        # 加入总列表,最后合并存总表
        all_records.extend(current_channel_records)
    time.sleep(REQUEST_INTERVAL)

# 合并存总表
if not SAVE_SEPARATE_CSV and all_records:
    total_df = pd.DataFrame(all_records)
    total_df.to_csv('all_channel_buy.csv', index=False, encoding='utf-8-sig')
    print(f"总数据量:{len(total_df)},已存储完成")
适配批量场景的优化说明
  • 完全保留名称和频道ID的一一对应关系,无需转换集合,支持上千组数据的批量处理
  • 自带请求重试和限流自动等待逻辑,避免批量请求时程序意外中断
  • 支持两种存储模式,可根据需求选择单频道单独存储或全量数据合并存储
  • 所有变量名做了区分,不会出现覆盖问题,返回数据异常时会自动跳过并打印日志,不会影响后续频道处理

内容的提问来源于stack exchange,提问作者joe8ones

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.23 19:15:03