基于自定义规则标记IP地址重复项的R语言实现问询
Great job getting the full IP duplication working! Let's extend that code to handle your two partial duplication scenarios. We'll use string manipulation to create custom grouping keys that match your patterns, then flag all entries sharing those keys.
Scenario 2: Flag IPs with Matching First Three Segments (100.100.10.*** )
We want to mark every IP that shares the same first three octets with at least one other IP in the dataset.
Solution Code
require(dplyr) require(stringr) # Your original dataset data <- data.frame(ip = c("100.100.10.100","100.100.10.101", "100.100.10.102", "100.100.10.113", "100.100.10.100", "100.100.10.200", "100.100.11.100", "100.100.10.101", "100.100.10.10", "100.100.10.1")) # Extract first three segments and flag duplicates data <- data %>% mutate( ip_prefix = str_extract(ip, "^\\d+\\.\\d+\\.\\d+"), # Grab up to the third dot ip.dupe = ifelse(duplicated(ip_prefix) | duplicated(ip_prefix, fromLast = TRUE), 1, 0) )
Expected Output
| ip | ip_prefix | ip.dupe |
|---|---|---|
| 100.100.10.100 | 100.100.10 | 1 |
| 100.100.10.101 | 100.100.10 | 1 |
| 100.100.10.102 | 100.100.10 | 1 |
| 100.100.10.113 | 100.100.10 | 1 |
| 100.100.10.100 | 100.100.10 | 1 |
| 100.100.10.200 | 100.100.10 | 1 |
| 100.100.11.100 | 100.100.11 | 0 |
| 100.100.10.101 | 100.100.10 | 1 |
| 100.100.10.10 | 100.100.10 | 1 |
| 100.100.10.1 | 100.100.10 | 1 |
Logic breakdown: The regex ^\\d+\\.\\d+\\.\\d+ captures everything from the start of the IP up to the third dot. We then use your existing duplication check (duplicated() + fromLast) but apply it to this prefix instead of the full IP.
Scenario 3: Flag IPs with Matching First Three Segments + First Digit of 3-Digit Last Segment (100.100.10.1** )
Based on your expected output, this scenario focuses on IPs where the fourth octet is 3 digits (matching your 100.100.10.1** example). We'll mark these IPs if they share the first three segments and the first digit of their 3-digit fourth octet.
Solution Code
require(dplyr) require(stringr) data <- data.frame(ip = c("100.100.10.100","100.100.10.101", "100.100.10.102", "100.100.10.113", "100.100.10.100", "100.100.10.200", "100.100.11.100", "100.100.10.101", "100.100.10.10", "100.100.10.1")) data <- data %>% mutate( # Extract the fourth octet to check its length octet4 = str_split_fixed(ip, "\\.", 4)[,4], # Create group key: for 3-digit octets, use prefix + first digit; else unique key ip_group_key = case_when( str_length(octet4) == 3 ~ paste0(str_extract(ip, "^\\d+\\.\\d+\\.\\d+"), ".", str_sub(octet4, 1, 1)), TRUE ~ paste0(ip, "_unique") # Ensure non-3-digit octets don't get flagged ), # Flag duplicates based on the group key ip.dupe = ifelse(duplicated(ip_group_key) | duplicated(ip_group_key, fromLast = TRUE), 1, 0) ) %>% select(-octet4) # Optional: remove helper column to clean up output
Expected Output
| ip | ip_group_key | ip.dupe |
|---|---|---|
| 100.100.10.100 | 100.100.10.1 | 1 |
| 100.100.10.101 | 100.100.10.1 | 1 |
| 100.100.10.102 | 100.100.10.1 | 1 |
| 100.100.10.113 | 100.100.10.1 | 1 |
| 100.100.10.100 | 100.100.10.1 | 1 |
| 100.100.10.200 | 100.100.10.2 | 0 |
| 100.100.11.100 | 100.100.11.1 | 0 |
| 100.100.10.101 | 100.100.10.1 | 1 |
| 100.100.10.10 | 100.100.10.10_unique | 0 |
| 100.100.10.1 | 100.100.10.1_unique | 0 |
Logic breakdown: We first split the IP to get the fourth octet. For 3-digit octets, we build a key from the first three segments plus the first digit of the fourth. For shorter octets, we assign a unique key so they won't be marked as duplicates. The duplication check then flags all entries sharing the same valid group key.
内容的提问来源于stack exchange,提问作者blazej

