You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用正则表达式将宽格式DataFrame部分列重塑为长格式

解决宽格式DataFrame转换为目标长格式的问题

我来帮你搞定这个格式转换的问题~你已经成功用gather()+separate()得到了拆分后的整洁数据,现在想要把Dir和Degree合并成DirDegree列(比如D.13),其实有两种简单的方法:

方法一:基于已有tidy数据合并列(最直观)

既然你已经通过separate()把key拆成了Dir、Degree、Type,直接用paste()合并前两列就好,步骤简单不容易出错:

library(tidyverse)

# 你已经得到的整洁数据
df.tidy = df %>% gather(key, Wh, -time) %>% separate(key, c("Dir", "Degree", "Type"), "\\.")

# 合并Dir和Degree为DirDegree,然后调整列顺序
df.target <- df.tidy %>%
  mutate(DirDegree = paste(Dir, Degree, sep = ".")) %>%
  select(time, DirDegree, Type, Wh)

# 查看结果
head(df.target)

运行后就能得到你期望的格式:

time DirDegree Type        Wh
1 2009-01-01     D.13   JA -2.2529442
2 2009-01-02     D.13   JA  1.0698570
3 2009-01-03     D.13   JA  1.7558374
4 2009-01-04     D.13   JA -0.2571767
5 2009-01-05     D.13   JA  0.2389923
6 2009-01-06     D.13   JA  1.6444592

方法二:用extract()一步到位(修正正则表达式)

你之前尝试extract()时得到NA,是因为正则表达式写得不对。你的key格式是X.Y.Z(比如D.13.JA),我们需要把X.Y作为第一组(DirDegree),Z作为第二组(Type),正确的正则应该用[^.]+匹配任意非点的字符,确保适配不同长度的Dir或Degree:

df.target <- df %>%
  gather(key, Wh, -time) %>%
  extract(key, c("DirDegree", "Type"), "(^[^.]+\\.[^.]+)\\.([^.]+)$")

# 查看结果
head(df.target)

正则表达式解释:

  • ^[^.]+\\.[^.]+:匹配从开头到第二个点之前的所有字符(比如D.13),其中[^.]+表示任意非点的字符序列,\\.匹配点本身
  • \\.:匹配分隔DirDegree和Type的点
  • ([^.]+)$:匹配最后一个点之后的所有字符(比如JA),$表示字符串结尾,确保匹配完整的Type部分

这样就能正确提取出你需要的列,不会出现NA啦~

内容的提问来源于stack exchange,提问作者tRash

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 07:40:39