You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言Dataframe重塑:按区块、孔位整理测试及浓度数据

问题与代码扩展性分析

原始数据与需求

输入数据框

df <- structure(list(block = c(1, 1, 1, 1, 1, 2, 2, 2, 2, 2), Well = c("A01", "A02", "A03", "A04", "A05", "A01", "A02", "A03", "A04", "A05"), Conc.test1 = c(2, NA, 2, NA, NA, 2, NA, 2, NA, NA), Conc.test2 = c(NA, 2, NA, 2, NA, NA, 2, NA, 2, NA), Conc.Test3 = c(NA, NA, NA, 2, NA, NA, NA, NA, 2, NA), Conc.test4 = c(NA, NA, 2, 2, NA, 2, NA, 2, 2, NA), Conc.test5 = c(NA, 2, NA, NA, NA, NA, 2, NA, NA, NA), Conc.test6 = c(2, NA, NA, NA, NA, 2, NA, NA, NA, NA)), row.names = c(NA, 10L), class = "data.frame")

#    block Well Conc.test1 Conc.test2 Conc.Test3 Conc.test4 Conc.test5 Conc.test6
# 1      1  A01          2         NA         NA         NA         NA          2
# 2      1  A02         NA          2         NA         NA          2         NA
# 3      1  A03          2         NA         NA          2         NA         NA
# 4      1  A04         NA          2          2          2         NA         NA
# 5      1  A05         NA         NA         NA         NA         NA         NA
# 6      2  A01          2         NA         NA          2         NA          2
# 7      2  A02         NA          2         NA         NA          2         NA
# 8      2  A03          2         NA         NA          2         NA         NA
# 9      2  A04         NA          2          2          2         NA         NA
# 10     2  A05         NA         NA         NA         NA         NA         NA

输出需求

按区块(block)、孔位(Well)统计所用测试的数量、测试名称及对应浓度,输出格式如下:

structure(list(block = c(1, 1, 1, 1, 2, 2, 2, 2), tests = c(2, 2, 2, 3, 3, 2, 2, 3), well = c("A01", "A02", "A03", "A04", "A01", "A02", "A03", "A04"), C1 = c("test1", "test2", "test1", "test2", "test1", "test2", "test1", "test2"), conc1 = c(2, 2, 2, 2, 2, 2, 2, 2), C2 = c("test6", "test5", "test4", "test3", "test6", "test5", "test4", "test3"), con2 = c(2, 2, 2, 2, 2, 2, 2, 2), C3 = c(NA, NA, NA, "test4", "test4", NA, NA, "test4"), conc3 = c(NA, NA, NA, 2, 2, NA, NA, 2)), row.names = c(NA, 8L), class = "data.frame")

#   block tests well    C1 conc1    C2 con2    C3 conc3
# 1     1     2  A01 test1     2 test6    2  <NA>    NA
# 2     1     2  A02 test2     2 test5    2  <NA>    NA
# 3     1     2  A03 test1     2 test4    2  <NA>    NA
# 4     1     3  A04 test2     2 test3    2 test4     2
# 5     2     3  A01 test1     2 test6    2 test4     2
# 6     2     2  A02 test2     2 test5    2  <NA>    NA
# 7     2     2  A03 test1     2 test4    2  <NA>    NA
# 8     2     3  A04 test2     2 test3    2 test4     2

现有实现代码

output_df <- df %>%
  pivot_longer(cols = starts_with("Conc"), names_to = "Test", values_to = "Concentration") %>%
  filter(!is.na(Concentration)) %>%
  group_by(block, Well) %>%
  mutate(test_num = row_number()) %>%
  pivot_wider(names_from = test_num, values_from = c(Test, Concentration), 
              names_sep = "") %>%
  ungroup()

output_df <- output_df %>%
  mutate(tests = rowSums(!is.na(select(., starts_with("Test"))))) %>%
  select(block, Well, tests, everything()) %>%
  rename_with(~gsub("Test", "C", .x), starts_with("Test")) %>%
  rename_with(~gsub("Concentration", "conc", .x), starts_with("Concentration"))

扩展性分析与优化建议

扩展性确认

现有代码的扩展性非常优秀,具体体现在:

  • 适配新增测试列:通过starts_with("Conc")匹配所有测试列,后续新增任意以Conc.开头的测试列(如Conc.test7),代码无需修改即可自动识别处理。
  • 动态适配测试数量:利用row_number()为每个孔位的测试生成序号,pivot_wider会根据每个孔位的实际测试数量动态生成对应数量的Cn和conc n列,无需硬编码列数。
  • 自动统计测试数量:rowSums(!is.na(select(., starts_with("Test"))))会自动统计所有测试名称列的非NA值数量,新增测试列后仍能正确计算tests字段。
  • 通用列名重命名:rename_with结合正则表达式批量修改列名,无论生成多少测试相关列,都能自动将Test{n}转为C{n}、Concentration{n}转为conc{n}。

简化优化(可选)

可以将两步管道合并为一个,减少中间变量,让代码更紧凑:

output_df <- df %>%
  pivot_longer(cols = starts_with("Conc"), names_to = "Test", values_to = "Concentration") %>%
  filter(!is.na(Concentration)) %>%
  group_by(block, Well) %>%
  mutate(test_num = row_number()) %>%
  pivot_wider(names_from = test_num, values_from = c(Test, Concentration), 
              names_sep = "") %>%
  ungroup() %>%
  mutate(tests = rowSums(!is.na(select(., starts_with("Test"))))) %>%
  select(block, Well, tests, everything()) %>%
  rename_with(~gsub("Test", "C", .x), starts_with("Test")) %>%
  rename_with(~gsub("Concentration", "conc", .x), starts_with("Concentration"))

内容的提问来源于stack exchange,提问作者szmple

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 16:07:03