Python Pandas如何根据retweet_count列值重复对应tweet行数据
实现方案
你现有代码中不需要将retweet_count转为字符串,我们需要保留它的整数属性来计算每行的重复次数。
pandas内置的repeat方法可以直接实现按指定次数重复行的需求,你只需要给该方法传入每行对应的重复次数(即retweet_count + 1)即可,之后过滤保留需要的created_at和tweet两列就符合输出要求。
完整可运行代码如下:
import pandas as pd def iterateTweets(): # 读取csv文件 tweets = pd.read_csv(r'data_set_tweets.csv') df = pd.DataFrame(tweets, columns=['created_at', 'tweet', 'retweet_count']) # 类型处理,注意保留retweet_count的整数类型 df['created_at'] = pd.to_datetime(df['created_at']) df['tweet'] = df['tweet'].astype(str) # 按retweet_count+1的次数重复行,reset_index用于重置连续的行索引 result_df = df.repeat(df['retweet_count'] + 1).reset_index(drop=True) # 仅保留需要的两列数据 result_df = result_df[['created_at', 'tweet']] return result_df if __name__ == '__main__': res = iterateTweets() # 如果需要导出为csv文件可以取消注释下行代码 # res.to_csv('output_result.csv', index=False) print(res)
如果需要输出的created_at和原数据格式完全一致,不需要转成datetime类型,删除df['created_at'] = pd.to_datetime(df['created_at'])这行即可。
内容的提问来源于stack exchange,提问作者Nick_Code
相关产品推荐
相关产品推荐

