You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用select(contains("T"))时time列丢失,pivot_longer异常排查

问题:pivot_longer搭配contains("T")时丢失time列

在R语言数据处理场景中,使用select()搭配contains("T")参数后,再用pivot_longer()会丢失time列,但使用contains("PAR")或contains("RH")时函数运行正常;改用gather()则无此异常。

数据子集

# ad_1_top_n <- ad_1[1:3, ]
# dput(ad_1_top_n)

ad_1 <- structure(list(time = 0:2, X1c.PAR1 = c(125.73, 125.76, 125.73
), X1c.PAR2 = c(27.25, 27.25, 27.25), X1c.PAR3 = c(10.03, 10.14, 
9.99), X1c.PAR4 = c(0.07, 0.07, 0.07), X1c.T1 = c(21.28, 21.28, 
21.27), X1c.T2 = c(20.19, 20.18, 20.18), X1c.T3 = c(19.95, 19.95, 
19.95), X1c.T4 = c(20.42, 20.4, 20.4), X1c.RH1 = c(73.57, 73.67, 
73.83), X1c.RH2 = c(84.02, 84.04, 84.06), X1c.RH3 = c(82.35, 
82.35, 82.36), X1c.RH4 = c(76.94, 76.91, 76.94), X1b.PAR1 = c(108.89, 
108.93, 108.93), X1b.PAR2 = c(61.9, 61.9, 61.9), X1b.PAR3 = c(40.1, 
40.1, 40.1), X1b.PAR4 = c(0.06, 0.06, 0.1), X1b.T1 = c(20.45, 
20.45, 20.46), X1b.T2 = c(20.22, 20.26, 20.26), X1b.T3 = c(20.63, 
20.6, 20.61), X1b.T4 = c(20.9, 20.89, 20.89), X1b.RH1 = c(76.2, 
76.2, 76.26), X1b.RH2 = c(79.18, 79.21, 79.21), X1b.RH3 = c(75.33, 
75.36, 75.47), X1b.RH4 = c(73.52, 73.58, 73.64), X1p.PAR1 = c(144.9, 
144.97, 144.94), X1p.PAR2 = c(116.96, 117.03, 116.96), X1p.PAR3 = c(85.2, 
85.24, 85.2), X1p.PAR4 = c(2.73, 2.73, 2.73), X1p.T1 = c(20.49, 
20.5, 20.48), X1p.T2 = c(20.62, 20.62, 20.61), X1p.T3 = c(19.99, 
20, 19.98), X1p.T4 = c(20.31, 20.3, 20.27), X1p.RH1 = c(77.79, 
77.87, 77.97), X1p.RH2 = c(77.01, 77.04, 77.05), X1p.RH3 = c(81.93, 
82.03, 82.08), X1p.RH4 = c(77.98, 77.75, 77.72), X2c.PAR1 = c(NaN, 
NaN, NaN), X2c.PAR2 = c(NaN, NaN, NaN), X2c.PAR3 = c(NaN, NaN, 
NaN), X2c.PAR4 = c(NaN, NaN, NaN), X2c.T1 = c(NaN, NaN, NaN), 
    X2c.T2 = c(NaN, NaN, NaN), X2c.T3 = c(NaN, NaN, NaN), X2c.T4 = c(NaN, 
    NaN, NaN), X2c.RH1 = c(NaN, NaN, NaN), X2c.RH2 = c(NaN, NaN, 
    NaN), X2c.RH3 = c(NaN, NaN, NaN), X2c.RH4 = c(NaN, NaN, NaN
    ), X2b.PAR1 = c(NaN, NaN, NaN), X2b.PAR2 = c(NaN, NaN, NaN
    ), X2b.PAR3 = c(NaN, NaN, NaN), X2b.PAR4 = c(NaN, NaN, NaN
    ), X2b.T1 = c(NaN, NaN, NaN), X2b.T2 = c(NaN, NaN, NaN), 
    X2b.T3 = c(NaN, NaN, NaN), X2b.T4 = c(NaN, NaN, NaN), X2b.RH1 = c(NaN, 
    NaN, NaN), X2b.RH2 = c(NaN, NaN, NaN), X2b.RH3 = c(NaN, NaN, 
    NaN), X2b.RH4 = c(NaN, NaN, NaN), X2p.PAR1 = c(NaN, NaN, 
    NaN), X2p.PAR2 = c(NaN, NaN, NaN), X2p.PAR3 = c(NaN, NaN, 
    NaN), X2p.PAR4 = c(NaN, NaN, NaN), X2p.T1 = c(NaN, NaN, NaN
    ), X2p.T2 = c(NaN, NaN, NaN), X2p.T3 = c(NaN, NaN, NaN), 
    X2p.T4 = c(NaN, NaN, NaN), X2p.RH1 = c(NaN, NaN, NaN), X2p.RH2 = c(NaN, 
    NaN, NaN), X2p.RH3 = c(NaN, NaN, NaN), X2p.RH4 = c(NaN, NaN, 
    NaN)), row.names = c(NA, 3L), class = "data.frame")

出现问题的处理代码

第三个处理函数会丢失time列:

library(tidyr)

adl_1_PAR_pivot <- ad_1 %>% select(time, contains("PAR")) %>% 
  pivot_longer(cols = contains("PAR"), names_to = "key", values_to = "PAR") %>% 
separate(col = "key", into = c("Dummy1", "Location", "Position", "Dummy2", "Height"), 
         sep = c(1,2,3,7), remove = TRUE) %>%
select(-Dummy1, -Dummy2)

adl_1_RH_pivot <- ad_1 %>% select(time, contains("RH")) %>% 
  pivot_longer(cols = contains("RH"), names_to = "key", values_to = "RH") %>% 
  separate(col = "key", into = c("Dummy1", "Location", "Position", "Dummy2", "Height"), 
           sep = c(1,2,3,6), remove = TRUE) %>%
  select(-Dummy1, -Dummy2)

adl_1_T_pivot <- ad_1 %>% select(time, contains("T")) %>% 
  pivot_longer(cols = contains("T"), names_to = "key", values_to = "T") %>% 
  separate(col = "key", into = c("Dummy1", "Location", "Position", "Dummy2", "Height"), 
           sep = c(1,2,3,5), remove = TRUE) %>%
  select(-Dummy1, -Dummy2)

正常运行的gather代码

adl_1_T <- ad_1 %>% select(time, contains("T")) %>% 
  gather(key = key, value = T, -time) %>%
  separate(col = "key", into = c("Dummy1","Location", "Position", "Dummy2","Height"), 
           sep = c(1,2,3,5), remove = TRUE) %>%
  select(-Dummy1, -Dummy2)

原因分析与解决方案

原因

contains("T")会匹配所有列名中包含字母"T"的列,而time列名里恰好有"T"。在pivot_longer的cols参数中使用contains("T")时,time列被误选为需要宽转长的列,最终导致time列被转换后丢失。

而使用contains("PAR")或contains("RH")时,time列名不包含这些字符串,所以不会被误选;gather()中用-time明确排除了time列,因此也不会出现问题。

解决方法

方法1:明确排除time列

在pivot_longer的cols参数中用-time指定要转换的列:

adl_1_T_pivot <- ad_1 %>% select(time, contains("T")) %>% 
  pivot_longer(cols = -time, names_to = "key", values_to = "T") %>% 
  separate(col = "key", into = c("Dummy1", "Location", "Position", "Dummy2", "Height"), 
           sep = c(1,2,3,5), remove = TRUE) %>%
  select(-Dummy1, -Dummy2)

方法2:更精确匹配目标列

使用matches("\\.T")匹配列名中包含.T的列(避免匹配time):

adl_1_T_pivot <- ad_1 %>% select(time, matches("\\.T")) %>% 
  pivot_longer(cols = matches("\\.T"), names_to = "key", values_to = "T") %>% 
  separate(col = "key", into = c("Dummy1", "Location", "Position", "Dummy2", "Height"), 
           sep = c(1,2,3,5), remove = TRUE) %>%
  select(-Dummy1, -Dummy2)

内容的提问来源于stack exchange,提问作者ivo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 22:25:12