You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将字典列表指定键值加载为pandas DataFrame

从JSON字符串列表提取指定字段生成pandas DataFrame

你可以通过JSON解析+字段提取的方式实现需求,不需要额外依赖第三方工具,以下是可直接运行的实现方案:

方法1:逐行解析(兼容性最佳)

逐行解析JSON字符串后手动提取目标字段,适配部分推文数据字段缺失的场景:

import json
import pandas as pd

# 原始推文列表
tweet_list = ['{"created_at": "Tue Aug 01 00:17:27 +0000 2017", "id": 892177421306343426, "id_str": "892177421306343426", "full_text": "This is Tilly. She\'s just checking pup on you.", "truncated": false, "display_text_range": [0, 138], "contributors": null, "is_quote_status": false, "retweet_count": 6514, "favorite_count": 33819, "favorited": false, "retweeted": false, "possibly_sensitive": false, "possibly_sensitive_appealable": false, "lang": "en"}',
 '{"created_at": "Sun Jul 30 15:58:51 +0000 2017", "id": 891689557279858688, "id_str": "891689557279858688", "full_text": "This is Darla. She commenced a snooze mid meal.", "truncated": false, "display_text_range": [0, 79], "entities": {"hashtags": [], "symbols": [],  "following": true, "follow_request_sent": false, "notifications": false, "translator_type": "none"}, "geo": null, "coordinates": null, "place": null, "contributors": null, "is_quote_status": false, "retweet_count": 8964, "favorite_count": 42908, "favorited": false, "retweeted": false, "possibly_sensitive": false, "possibly_sensitive_appealable": false, "lang": "en"}']

result = []
for item in tweet_list:
    # 把JSON格式字符串转为Python字典
    tweet = json.loads(item)
    # 提取需要的三个字段,将原id键对应值映射到ID列
    result.append({
        "ID": tweet["id"],
        "retweet_count": tweet["retweet_count"],
        "favorite_count": tweet["favorite_count"]
    })

df = pd.DataFrame(result)

运行后得到的DataFrame内容如下:

IDretweet_countfavorite_count
892177421306343426651433819
891689557279858688896442908

方法2:pandas内置方法(代码更简洁)

如果所有推文JSON结构一致,没有字段缺失的问题,可以直接用pandas内置的JSON读取接口减少代码量:

import pandas as pd

# 把列表转为换行分隔的JSON lines格式,直接读取为DataFrame
df = pd.read_json("\n".join(tweet_list), lines=True)
# 筛选目标列,将id列重命名为ID
df = df[["id", "retweet_count", "favorite_count"]].rename(columns={"id": "ID"})

提示:如果存在部分推文缺少目标字段的情况,建议使用方法1,提取字段时可以用tweet.get("id", None)的写法给缺失值填充默认空值,避免程序报错中断。

内容的提问来源于stack exchange,提问作者Nicole Sai

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 23:24:14