You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python3.4 Twitter数据提取报错:KeyError: 'user' 问题咨询

问题分析与解决方案

核心结论

是的,你采集的推文完全有可能缺失user字段,这就是触发KeyError: 'user'的根本原因。Twitter的JSON数据里并非所有条目都是标准推文对象——比如部分系统通知、特殊转发条目、采集过程中损坏的记录,甚至API返回的非推文类型数据,都可能没有user键,哪怕你直觉上认为用户ID“不可能为空”。

错误原因

你的代码直接通过tweet['user']['id']访问嵌套字段,但没有先确认tweet中是否存在user这个顶级键。一旦遇到不含user的条目,Python就会抛出KeyError中断程序运行。

修复后的代码

我们可以通过先检查键是否存在,或者使用字典的get()方法安全访问嵌套字段,彻底避免这类错误。下面是优化后的完整代码:

import json

fname = 'test_with_sample.json'
with open(fname, 'r') as f:
    users_with_geodata = {
        "data": []
    }
    all_users = []
    total_tweets = 0

    for line in f:
        # 先处理格式损坏的JSON行
        try:
            tweet = json.loads(line)
        except json.JSONDecodeError:
            continue
        
        # 安全检查:确保user和user.id字段存在
        if 'user' not in tweet or 'id' not in tweet['user']:
            continue
        
        total_tweets += 1
        user_id = tweet['user']['id']
        
        if user_id not in all_users:
            all_users.append(user_id)
            # 用get()方法处理可能缺失的字段,避免后续报错
            user_data = {
                "user_id": user_id,
                "features": {
                    "name": tweet['user'].get('name', ''),
                    "id": user_id,
                    "screen_name": tweet['user'].get('screen_name', ''),
                    "tweets": 1,
                    "location": tweet['user'].get('location', ''),
                }
            }
            
            # 处理推文地理位置
            if tweet.get('place'):
                user_data["features"]["primary_geo"] = f"{tweet['place'].get('full_name', '')}, {tweet['place'].get('country', '')}"
                user_data["features"]["geo_type"] = "Tweet place"
            else:
                user_data["features"]["primary_geo"] = tweet['user'].get('location', '')
                user_data["features"]["geo_type"] = "User location"
            
            # 仅保留有地理数据的用户
            if user_data["features"]["primary_geo"]:
                users_with_geodata['data'].append(user_data)
        else:
            # 更新已有用户的推文计数,找到后立即退出循环提升效率
            for user in users_with_geodata["data"]:
                if user_id == user["user_id"]:
                    user["features"]["tweets"] += 1
                    break

    # 计算带地理数据的总推文数
    total_geo_tweets = sum(user["features"]["tweets"] for user in users_with_geodata["data"])

    # 输出统计结果
    print(f"The file included {len(all_users)} unique users who tweeted with or without geo data")
    print(f"The file included {len(users_with_geodata['data'])} unique users who tweeted with geo data, including 'location'")
    print(f"The users with geo data tweeted {total_geo_tweets} out of the total {total_tweets} of tweets.")

# 保存结果到JSON文件
with open('users_geo_sample.json', 'w') as fout:
    json.dump(users_with_geodata, fout, indent=4)

关键优化点

  • 新增try-except捕获JSON格式错误的行,避免因数据损坏中断程序
  • 增加'user' in tweet和'id' in tweet['user']的前置检查,跳过无用户信息的无效条目
  • 使用字典get()方法访问字段,字段缺失时返回默认空字符串,避免后续嵌套访问报错
  • 修正原代码中geo_tweets重复累加的逻辑错误
  • 用f-string替代字符串拼接,提升代码可读性与简洁度

内容的提问来源于stack exchange,提问作者Bonzay

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 07:04:33