You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何对数据框每一行应用approx插值函数进行转换?

对数据框每行应用线性插值的实现方法

问题背景

现有一个数据框,每行对应不同的儿童照料类型,列是特定月龄(0、6、15、24、36个月)的占比数据。需要对每行使用approx函数进行线性插值,得到0到36个月所有月龄的对应值,尝试用rowwise()结合group_modify时失败,需找到正确实现方式。

数据预处理(沿用你的代码)

library(tidyverse)
library(stats)

care_distribution_table = "Age,M0,M6,M15,M24,M36
Mother,100.00%,36.89%,29.58%,29.61%,26.64%
Father,0.00%,11.74%,14.70%,12.23%,11.40%
Other relative,0.00%,15.80%,14.00%,11.55%,10.20%
In-home nonrelative,0.00%,7.94%,8.89%,7.33%,6.48%
Home-based childcare,0.00%,18.62%,20.69%,21.82%,17.38%
Center childcare,0.00%,9.00%,12.15%,17.46%,27.90%"

care_dist = read.csv(text = care_distribution_table, row.names = 1) |>
  # 把百分比转为小数
  mutate(across(everything(), ~as.numeric(trimws(.x, whitespace = '%')) / 100)) |>
  # 修正四舍五入误差,确保每列和为1
  mutate(across(everything(), ~.x/sum(.x))) 

解决方案1:rowwise() + mutate存储插值结果并展开(长格式)

先将每行插值结果存为列表列,再拆分为标准长格式,方便后续可视化或分析:

interpolated_df = care_dist |>
  rowwise() |>
  # 对当前行所有列执行插值,结果存为列表
  mutate(interpolated = list(approx(x = c(0,6,15,24,36), y = c_across(everything()), xout = 0:36))) |>
  ungroup() |>
  # 提取插值的月龄和对应值,展开为单独行
  mutate(age = list(interpolated$x), value = list(interpolated$y)) |>
  select(-interpolated) |>
  unnest(c(age, value)) |>
  # 将原行名转为标识列
  mutate(care_type = rownames(care_dist)) |>
  relocate(care_type, age, value)

# 查看前几行结果
head(interpolated_df)

解决方案2:修正group_modify用法(你的失败代码修复)

你之前的代码失败是因为误用了不存在的df函数,改用tibble构建结果即可:

# 先将原行名转为分组列
care_dist_with_id = care_dist |>
  rownames_to_column(var = "care_type") |>
  group_by(care_type)

# 对每个分组(每行)执行插值
interpolated_result = care_dist_with_id |>
  group_modify(function(.x, .y) {
    # 提取当前行的数值(排除care_type列)
    y_vals = .x |> select(-care_type) |> unlist()
    # 执行插值
    interp = approx(x = c(0,6,15,24,36), y = y_vals, xout = 0:36)
    # 返回标准tibble格式结果
    tibble(age = interp$x, value = interp$y)
  })

# 查看前几行结果
head(interpolated_result)

解决方案3:基础R的apply逐行处理(宽格式输出)

如果偏好基础R语法,可直接用apply逐行插值,输出宽格式数据框:

# 逐行执行插值,结果转置为矩阵
interp_matrix = t(apply(care_dist, 1, function(row) {
  approx(x = c(0,6,15,24,36), y = row, xout = 0:36)$y
}))

# 转换为数据框并设置列名、行名
interp_df = as.data.frame(interp_matrix)
colnames(interp_df) = paste0("M", 0:36)
rownames(interp_df) = rownames(care_dist)

# 查看前5列(月龄0-4)的结果
head(interp_df[,1:5])

内容的提问来源于stack exchange,提问作者Mohan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 10:15:59