求助:使用多分隔符拆分DataFrame列并生成目标格式数据
No problem! When dealing with multiple delimiters in a column that needs to be split into rows while repeating the corresponding ID (here ContractCd), there are straightforward ways to solve this in R—either using the tidyverse for simplicity or base R if you prefer avoiding extra packages.
Sample Data
First, let's recap your input data for clarity:
ContractCd <- c(9940099251,9940080497,9940099251,9940014221) WBSElementNbr <- c("N1075001,N1075013,MT842001,N1128001,NN480001,N1142001,N1147001","IV768001&IU775001","NN480001;N1147001","D6268001/D6268005") Data <- data.frame(ContractCd,WBSElementNbr)
Solution 1: Using tidyverse (Recommended for Simplicity)
The separate_rows() function from tidyr (part of the tidyverse suite) supports regular expressions to handle multiple delimiters. Just specify all your delimiters in a regex character class:
# Load the tidyverse package library(tidyverse) # Split the column and expand rows Data1 <- Data %>% separate_rows(WBSElementNbr, sep = "[,;&/]")
The regex [,;&/] tells R to split on any of the characters inside the brackets: comma, semicolon, ampersand, or slash. This will generate exactly the expanded output you're expecting.
Solution 2: Base R (No External Packages)
If you don't want to rely on external packages, you can achieve the same result with base R functions:
# Split each string into a list of individual elements split_wbs <- strsplit(Data$WBSElementNbr, "[,;&/]") # Create the expanded data frame Data1_base <- data.frame( ContractCd = rep(Data$ContractCd, sapply(split_wbs, length)), WBSElementNbr = unlist(split_wbs) )
Here, strsplit() splits each entry using the same regex, sapply() calculates how many times each ContractCd needs to be repeated, and unlist() flattens the split elements into a single vector for the new column.
Either approach will produce your desired Data1 structure, where each WBSElementNbr entry sits on its own row paired with the matching ContractCd.
内容的提问来源于stack exchange,提问作者JK1185

