Julia按组合并含NaN的行并去除空值
在Julia DataFrame中按组合并去除NaN值
原始数据
首先我们有如下结构的DataFrame:
using DataFrames df = DataFrame( group = ["A", "A", "A", "B", "B", "B"], V1 = [1, NaN, NaN, 3, NaN, NaN], V2 = [NaN, 4, NaN, 2, NaN, NaN], V3 = [NaN, NaN, 4, NaN, 1, NaN])
其输出为:
6×4 DataFrame Row │ group V1 V2 V3 │ String Float64 Float64 Float64 ─────┼─────────────────────────────────── 1 │ A 1.0 NaN NaN 2 │ A NaN 4.0 NaN 3 │ A NaN NaN 4.0 4 │ B 3.0 2.0 NaN 5 │ B NaN NaN 1.0 6 │ B NaN NaN NaN
实现方法
我们需要按group列分组,将每组内的行合并,保留各列的非NaN有效值,最终得到每组一行的结果。可以通过groupby结合combine函数实现:
方法1:指定列处理
直接针对每个数值列,提取分组内的非NaN有效值:
# 分组并聚合 result = combine(groupby(df, :group), :V1 => x -> first(skipmissing(x)), :V2 => x -> first(skipmissing(x)), :V3 => x -> first(skipmissing(x))) # 重命名列以恢复原名称 rename!(result, [:group, :V1, :V2, :V3])
方法2:批量处理所有数值列
如果列数较多,可以用批量方式处理所有数值类型的列,避免逐个指定:
result = combine(groupby(df, :group)) do gdf # 遍历所有非分组列,提取非NaN有效值 cols = [] for col in eachcol(gdf, false) col_name = nameof(col) if col_name == :group push!(cols, col_name => first(col)) else push!(cols, col_name => first(skipmissing(col))) end end DataFrame(cols) end
最终结果
执行上述代码后,会得到期望的输出:
2×4 DataFrame Row │ group V1 V2 V3 │ String Int64 Int64 Int64 ─────┼───────────────────────────── 1 │ A 1 4 4 2 │ B 3 2 1
说明
skipmissing(x)用于过滤列中的NaN值,first()提取过滤后的第一个有效值(因为每组内每个列仅存在一个非NaN值)。- 如果分组内某列存在多个非NaN值,可根据需求替换
first()为collect()、mean()等其他聚合函数。
内容的提问来源于stack exchange,提问作者Quinten
相关产品推荐
相关产品推荐

