如何将Tweepy Paginator获取的推文计数转换为Pandas DataFrame?
解决Tweepy推文计数转Pandas DataFrame的格式问题
你当前的问题出在嵌套列表的处理上:tweepy.Paginator返回的每页数据count.data是一个字典列表,而你用append()把整个列表作为元素加入tweet_count,最终得到了列表套列表的结构。只需要调整数据收集的方式,再转成DataFrame即可。
完整修正代码
import tweepy import pandas as pd import time # 假设你已初始化好client、query、start_time、end_time counts = tweepy.Paginator( client.get_all_tweets_count, query=query, start_time=start_time, end_time=end_time, granularity='day' ) time.sleep(1) tweet_count = [] # 用extend代替append,展开每页的字典列表 for page in counts: tweet_count.extend(page.data) # 转成DataFrame并调整列名与顺序 df = pd.DataFrame(tweet_count) df = df.rename(columns={ 'start': 'Start', 'end': 'End', 'tweet_count': 'Count' }) df = df[['Start', 'End', 'Count']] # 查看转换结果 print(df.head())
关键说明
- 替换append为extend:
append()会把整个page.data列表作为单个元素加入tweet_count,而extend()会把列表里的每个字典单独添加,最终得到单层字典列表——这是Pandas能直接解析的格式。 - 重命名列名:通过
rename()把原始字段名start/end/tweet_count改成你需要的Start/End/Count。 - 调整列顺序:用
df[['Start', 'End', 'Count']]确保列的排列和目标结构一致。
内容的提问来源于stack exchange,提问作者Laurence Bach
相关产品推荐
相关产品推荐

