Julia读取gzip压缩CSV文件并实现数据统计可视化求助
Julia处理gzip压缩CSV:日期提取、统计Top10及select函数问题解决
一、解决文件读取与select函数异常问题
你的代码同时加载了CSV和CSVFiles两个包,二者API存在冲突,会导致DataFrames的select函数无法正常调用。只需保留CSV包即可(它与DataFrames生态适配更好,且原生支持gzip压缩文件读取)。
简化后的文件读取代码:
import Pkg # 仅安装必要包 # Pkg.add(["CSV", "DataFrames", "Dates", "Plots"]) using CSV, DataFrames, Dates, Plots # CSV.read原生支持gzip后缀的文件,无需额外解压操作 df = CSV.read("Path//to//file//file.csv.gzip", DataFrame) # 现在可以正常使用select函数 selected_df = select(df, :thread_created_utc, :comment_created_utc, :thread_id, :comment_id, :user_id)
二、日期格式转换(对应Pandas代码逻辑)
根据thread_created_utc的类型,分两种情况处理:
情况1:字段为Unix时间戳(数值类型)
# 将Unix时间戳转为Date类型 df[!, :threadcreateddate] = Date.(unix2datetime.(df.thread_created_utc)) df[!, :commentcreateddate] = Date.(unix2datetime.(df.comment_created_utc))
情况2:字段为字符串格式的UTC时间
# 解析字符串为DateTime后转Date df[!, :threadcreateddate] = Date.(DateTime.(df.thread_created_utc, dateformat"yyyy-mm-ddTHH:MM:SSZ")) df[!, :commentcreateddate] = Date.(DateTime.(df.comment_created_utc, dateformat"yyyy-mm-ddTHH:MM:SSZ"))
注意:如果你的时间字符串格式不同,需要调整
dateformat参数匹配实际格式。
三、统计发帖量Top10日期
# 按日期分组,统计每日唯一发帖数 df_thread_counts = combine( groupby(df, :threadcreateddate), :thread_id => nunique => :thread_count ) # 排序并取Top10 top10_thread_dates = sort(df_thread_counts, :thread_count, rev=true)[1:10, :] println("发帖量Top10日期:") println(top10_thread_dates)
四、统计评论量Top10用户
# 按用户分组,统计每个用户的唯一评论数 df_user_comment_counts = combine( groupby(df, :user_id), :comment_id => nunique => :comment_count ) # 排序并取Top10 top10_comment_users = sort(df_user_comment_counts, :comment_count, rev=true)[1:10, :] println("评论量Top10用户:") println(top10_comment_users)
五、发帖量趋势可视化
# 绘制折线图 plot( df_thread_counts.threadcreateddate, df_thread_counts.thread_count, xlabel="日期", ylabel="发帖量", title="每日发帖量趋势", legend=false, linewidth=2 ) display(current())
内容的提问来源于stack exchange,提问作者Robin
相关产品推荐
相关产品推荐

