如何对数据框每一行应用approx插值函数进行转换?
对数据框每行应用线性插值的实现方法
问题背景
现有一个数据框,每行对应不同的儿童照料类型,列是特定月龄(0、6、15、24、36个月)的占比数据。需要对每行使用approx函数进行线性插值,得到0到36个月所有月龄的对应值,尝试用rowwise()结合group_modify时失败,需找到正确实现方式。
数据预处理(沿用你的代码)
library(tidyverse) library(stats) care_distribution_table = "Age,M0,M6,M15,M24,M36 Mother,100.00%,36.89%,29.58%,29.61%,26.64% Father,0.00%,11.74%,14.70%,12.23%,11.40% Other relative,0.00%,15.80%,14.00%,11.55%,10.20% In-home nonrelative,0.00%,7.94%,8.89%,7.33%,6.48% Home-based childcare,0.00%,18.62%,20.69%,21.82%,17.38% Center childcare,0.00%,9.00%,12.15%,17.46%,27.90%" care_dist = read.csv(text = care_distribution_table, row.names = 1) |> # 把百分比转为小数 mutate(across(everything(), ~as.numeric(trimws(.x, whitespace = '%')) / 100)) |> # 修正四舍五入误差,确保每列和为1 mutate(across(everything(), ~.x/sum(.x)))
解决方案1:rowwise() + mutate存储插值结果并展开(长格式)
先将每行插值结果存为列表列,再拆分为标准长格式,方便后续可视化或分析:
interpolated_df = care_dist |> rowwise() |> # 对当前行所有列执行插值,结果存为列表 mutate(interpolated = list(approx(x = c(0,6,15,24,36), y = c_across(everything()), xout = 0:36))) |> ungroup() |> # 提取插值的月龄和对应值,展开为单独行 mutate(age = list(interpolated$x), value = list(interpolated$y)) |> select(-interpolated) |> unnest(c(age, value)) |> # 将原行名转为标识列 mutate(care_type = rownames(care_dist)) |> relocate(care_type, age, value) # 查看前几行结果 head(interpolated_df)
解决方案2:修正group_modify用法(你的失败代码修复)
你之前的代码失败是因为误用了不存在的df函数,改用tibble构建结果即可:
# 先将原行名转为分组列 care_dist_with_id = care_dist |> rownames_to_column(var = "care_type") |> group_by(care_type) # 对每个分组(每行)执行插值 interpolated_result = care_dist_with_id |> group_modify(function(.x, .y) { # 提取当前行的数值(排除care_type列) y_vals = .x |> select(-care_type) |> unlist() # 执行插值 interp = approx(x = c(0,6,15,24,36), y = y_vals, xout = 0:36) # 返回标准tibble格式结果 tibble(age = interp$x, value = interp$y) }) # 查看前几行结果 head(interpolated_result)
解决方案3:基础R的apply逐行处理(宽格式输出)
如果偏好基础R语法,可直接用apply逐行插值,输出宽格式数据框:
# 逐行执行插值,结果转置为矩阵 interp_matrix = t(apply(care_dist, 1, function(row) { approx(x = c(0,6,15,24,36), y = row, xout = 0:36)$y })) # 转换为数据框并设置列名、行名 interp_df = as.data.frame(interp_matrix) colnames(interp_df) = paste0("M", 0:36) rownames(interp_df) = rownames(care_dist) # 查看前5列(月龄0-4)的结果 head(interp_df[,1:5])
内容的提问来源于stack exchange,提问作者Mohan
相关产品推荐
相关产品推荐

