为数据行对分配组编号:指定组对设唯一编号的技术需求
Solution for Assigning Group IDs to Row Pairs/Specific Groups
Got it, let's break down how to solve this problem. Based on your requirements, we need to:
- Assign unique incrementing IDs to specific row pairs (like
c("bad","good")) - Force the specific trio
c('Veni',"vidi","Vici")to use the fixed ID 666 - Ignore or mark non-group rows as needed
I'll use R for the example since your sample uses R-style vectors, but the logic translates easily to other languages like Python.
Step 1: Set Up Sample Data
First, let's create a sample dataset that mirrors your scenario, including random non-group rows:
data <- data.frame( text = c("random1", "bad", "good", "random2", "Veni", "vidi", "Vici", "random3", "bad", "good") )
Step 2: Handle the Special Fixed-ID Trio
We'll first identify and mark the Veni/vidi/Vici trio to ensure it gets ID 666:
library(dplyr) library(zoo) # For filling NA values across rows data <- data %>% # Flag the start of the special trio mutate(is_special = ifelse(text == "Veni" & lead(text, 1) == "vidi" & lead(text, 2) == "Vici", TRUE, FALSE)) %>% # Propagate the flag to all three rows in the trio mutate(is_special = na.locf(is_special, fromLast = TRUE, na.rm = FALSE)) %>% # Assign fixed ID 666 to these rows mutate(group_id = ifelse(is_special, 666, NA))
Step 3: Assign Incrementing IDs to Standard Row Pairs
Next, we'll handle the bad/good pairs, assigning unique incrementing IDs to each occurrence:
data <- data %>% # Flag the start of each bad/good pair mutate(is_pair = ifelse(text == "bad" & lead(text, 1) == "good", TRUE, FALSE)) %>% # Assign incrementing IDs to the start of each pair mutate(group_id = ifelse(is_pair, cumsum(is_pair), group_id)) %>% # Propagate the ID to the second row of each pair mutate(group_id = na.locf(group_id, na.rm = FALSE)) %>% # Clean up temporary flag columns select(-is_special, -is_pair)
Step 4: View the Result
Running the code above gives us exactly what we need:
print(data) # text group_id # 1 random1 NA # 2 bad 1 # 3 good 1 # 4 random2 NA # 5 Veni 666 # 6 vidi 666 # 7 Vici 666 # 8 random3 NA # 9 bad 2 #10 good 2
Key Notes for Adaptation
- If your row pairs/groups are defined differently (e.g., non-consecutive rows, different value pairs), just adjust the
is_specialoris_pairlogic to match your criteria. - If you want to assign IDs to non-group rows instead of leaving them as NA, replace the NA values with a default (like 0) using
mutate(group_id = ifelse(is.na(group_id), 0, group_id)). - For Python users, you can achieve the same result using
pandaswithshift()(instead oflead()) andffill()(instead ofna.locf()).
内容的提问来源于stack exchange,提问作者Alexander
相关产品推荐
相关产品推荐

