R语言新手求助:如何按唯一标识符去除行内连续重复值
Hey there! Since you're new to R, let's break down how to solve this problem clearly. From your example, it looks like you want to collapse consecutive identical non-NA values (0s or 1s) into a single instance, keeping only the sequence of value changes (so a run of 1s becomes one 1, a run of 0s becomes one 0, etc.). Here are two straightforward methods to do this:
Method 1: Base R (No Extra Packages Needed)
First, let's assume your dataset is named patient_data, with the first column as the unique patient ID and columns 2-35 as the 0/1/NA values. We'll create a custom function to process each row, then apply it to the whole dataset:
# Define a function to collapse consecutive non-NA values collapse_consecutive <- function(row) { # Step 1: Extract only non-NA values from the row non_na_vals <- row[!is.na(row)] # Step 2: Keep values that differ from the previous one # (The first value is always kept; then we check differences) collapsed_vals <- non_na_vals[c(TRUE, diff(non_na_vals) != 0)] # Optional: Turn the vector into a space-separated string paste(collapsed_vals, collapse = " ") } # Apply the function to every row (excluding the first ID column) patient_data$collapsed_sequence <- apply(patient_data[, -1], 1, collapse_consecutive)
For your example patient A, this will take the sequence 0 1 0 1 1...1 0 0...0 1 1...1 NA NA, strip out the NAs, then collapse consecutive duplicates to give you 0 1 0 1 0 1.
Method 2: Tidyverse (Using dplyr + purrr)
If you prefer using the tidyverse ecosystem (which is great for data manipulation as a beginner), here's how to do it with those packages:
# Load required packages (install first if you haven't: install.packages("tidyverse")) library(dplyr) library(purrr) patient_data <- patient_data %>% rowwise() %>% # Process each row individually mutate( # Combine columns 2-35 into a vector, and remove NA values clean_vals = list(c_across(2:35) %>% discard(is.na)), # Collapse consecutive duplicates and format as a string collapsed_sequence = list(paste(clean_vals[c(TRUE, diff(clean_vals) != 0)], collapse = " ")) ) %>% ungroup() %>% # Exit row-wise processing select(-clean_vals) # Remove the intermediate column if you don't need it
Quick Notes for Beginners
- If you want the result as a numeric vector instead of a string, just remove the
paste(...)part and keep the collapsed vector directly. - Double-check your column indices: if your value columns start at column 2 and end at 35,
2:35is correct. If your dataset has a different structure, adjust this range.
Let me know if you hit any snags—happy to help you refine this further!
内容的提问来源于stack exchange,提问作者Jara de la Court

