You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬取Reddit数据时遭遇KeyError问题求助

解决Reddit爬取中的KeyError: "id"异常问题

你的核心问题是异常处理范围不对——原来的try-except包裹了整个帖子遍历循环,只要有一个帖子缺失字段,整个批次的处理就会中断。下面是修改后的代码,实现单个帖子缺失字段时留空并继续执行的逻辑:

import pandas as pd
import requests
import time

final = pd.DataFrame()
oldestpost = False
for i in range(100): 
    # 发起请求,获取更早的帖子
    res = requests.get("https://oauth.reddit.com/r/europe/hot", headers=headers, params={"limit": 100, "after": oldestpost})

    listdf = list()
    
    # 逐个处理帖子,单独捕获异常
    for post in res.json()['data']['children']:
        try:
            post_data = post['data']
            listdf.append({
                'subreddit': post_data['subreddit'],
                'title': post_data['title'],
                'selftext': post_data['selftext'],
                'upvote_ratio': post_data['upvote_ratio'],
                'ups': post_data['ups'],
                'downs': post_data['downs'],
                'score': post_data['score'],
                'url': post_data["url"],
                "num_commments": post_data["num_comments"],
                "date": post_data["created"],
                "id": f"{post['kind']}_{post_data['id']}"
            })
        except KeyError as e:
            print(f"帖子缺失字段 {e},已补空")
            # 生成补空的条目,所有缺失字段用空字符串填充
            fallback_entry = {
                'subreddit': post.get('data', {}).get('subreddit', ''),
                'title': post.get('data', {}).get('title', ''),
                'selftext': post.get('data', {}).get('selftext', ''),
                'upvote_ratio': post.get('data', {}).get('upvote_ratio', ''),
                'ups': post.get('data', {}).get('ups', ''),
                'downs': post.get('data', {}).get('downs', ''),
                'score': post.get('data', {}).get('score', ''),
                'url': post.get('data', {}).get('url', ''),
                "num_commments": post.get('data', {}).get('num_comments', ''),
                "date": post.get('data', {}).get('created', ''),
                "id": ""  # id缺失时留空
            }
            listdf.append(fallback_entry)
            continue

    # 直接用列表生成DataFrame,简化逻辑
    df1 = pd.DataFrame(listdf)
    
    # 避免空DataFrame导致索引错误
    if not df1.empty:
        oldestpost = df1.iloc[-1]["id"]
        final = pd.concat([final, df1], ignore_index=True)
    else:
        print("本次请求无有效帖子,跳过")
    
    time.sleep(2)

关键修改说明

  • 缩小异常处理范围:把try-except放到单个帖子的处理逻辑内,单个帖子出错不影响其他帖子的抓取
  • 字段安全获取:用dict.get()方法替代直接索引,字段缺失时自动返回默认值(空字符串)
  • 补空逻辑:捕获KeyError时生成补空条目,确保数据完整性,同时继续执行循环
  • 简化DataFrame创建:直接用pd.DataFrame(listdf)替代循环concat,提升代码效率
  • 空数据判断:添加df1为空的判断,避免iloc[-1]触发索引错误

内容的提问来源于stack exchange,提问作者Frandersch

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 19:22:27