You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中基于Excel重复表头的值创建新列并重构数据集

解决方案

方案1:使用tidyverse套件实现(推荐)

首先加载依赖包:

library(dplyr)
library(purrr)

处理代码如下:

# 定位每个奶牛数据块的起始行(第一列为id的行)
chunk_starts <- which(df[, 1] == "id")

# 定义单块数据处理函数
process_chunk <- function(start_row) {
  # 提取当前块的奶牛编号
  cow_id <- df[start_row, 2]
  # 提取当前块的有效数据行
  data_rows <- df[(start_row + 2):(start_row + 3), ]
  # 整理为标准格式
  res <- data.frame(
    id = data_rows[, 1],
    Cow = cow_id,
    date = data_rows[, 3],
    day = data_rows[, 4],
    intake1 = data_rows[, 5],
    intake2 = data_rows[, 6],
    eattime = data_rows[, 7],
    visits = data_rows[, 8],
    dmi = data_rows[, 9],
    vtime = data_rows[, 10],
    stringsAsFactors = FALSE
  )
  return(res)
}

# 批量处理所有块并合并,转换数值列类型
df_final <- map_dfr(chunk_starts, process_chunk) %>%
  mutate(across(c(id, date, day, intake1, intake2, eattime, visits, dmi, vtime), as.numeric))

方案2:使用基础R实现(无需额外安装包)

# 定位每个奶牛数据块的起始行
chunk_starts <- which(df[, 1] == "id")
df_final <- data.frame()
col_names <- c("id", "Cow", "date", "day", "intake1", "intake2", "eattime", "visits", "dmi", "vtime")

# 循环处理每个数据块
for (i in chunk_starts) {
  cow_id <- df[i, 2]
  data_rows <- (i + 2):(i + 3)
  temp_df <- cbind(df[data_rows, 1], cow_id, df[data_rows, 3:10])
  colnames(temp_df) <- col_names
  df_final <- rbind(df_final, temp_df)
}

# 将数值类列转换为数值格式
num_cols <- c("id", "date", "day", "intake1", "intake2", "eattime", "visits", "dmi", "vtime")
df_final[num_cols] <- lapply(df_final[num_cols], function(x) as.numeric(as.character(x)))

两种方案输出的df_final与你期望的df2格式完全一致。


内容的提问来源于stack exchange,提问作者Jacquelyn

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 12:39:04