You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言新手求助:如何按唯一标识符去除行内连续重复值

How to Collapse Consecutive 0/1 Values (Ignoring NAs) in R

Hey there! Since you're new to R, let's break down how to solve this problem clearly. From your example, it looks like you want to collapse consecutive identical non-NA values (0s or 1s) into a single instance, keeping only the sequence of value changes (so a run of 1s becomes one 1, a run of 0s becomes one 0, etc.). Here are two straightforward methods to do this:

Method 1: Base R (No Extra Packages Needed)

First, let's assume your dataset is named patient_data, with the first column as the unique patient ID and columns 2-35 as the 0/1/NA values. We'll create a custom function to process each row, then apply it to the whole dataset:

# Define a function to collapse consecutive non-NA values
collapse_consecutive <- function(row) {
  # Step 1: Extract only non-NA values from the row
  non_na_vals <- row[!is.na(row)]
  # Step 2: Keep values that differ from the previous one
  # (The first value is always kept; then we check differences)
  collapsed_vals <- non_na_vals[c(TRUE, diff(non_na_vals) != 0)]
  # Optional: Turn the vector into a space-separated string
  paste(collapsed_vals, collapse = " ")
}

# Apply the function to every row (excluding the first ID column)
patient_data$collapsed_sequence <- apply(patient_data[, -1], 1, collapse_consecutive)

For your example patient A, this will take the sequence 0 1 0 1 1...1 0 0...0 1 1...1 NA NA, strip out the NAs, then collapse consecutive duplicates to give you 0 1 0 1 0 1.

Method 2: Tidyverse (Using dplyr + purrr)

If you prefer using the tidyverse ecosystem (which is great for data manipulation as a beginner), here's how to do it with those packages:

# Load required packages (install first if you haven't: install.packages("tidyverse"))
library(dplyr)
library(purrr)

patient_data <- patient_data %>%
  rowwise() %>% # Process each row individually
  mutate(
    # Combine columns 2-35 into a vector, and remove NA values
    clean_vals = list(c_across(2:35) %>% discard(is.na)),
    # Collapse consecutive duplicates and format as a string
    collapsed_sequence = list(paste(clean_vals[c(TRUE, diff(clean_vals) != 0)], collapse = " "))
  ) %>%
  ungroup() %>% # Exit row-wise processing
  select(-clean_vals) # Remove the intermediate column if you don't need it

Quick Notes for Beginners

  • If you want the result as a numeric vector instead of a string, just remove the paste(...) part and keep the collapsed vector directly.
  • Double-check your column indices: if your value columns start at column 2 and end at 35, 2:35 is correct. If your dataset has a different structure, adjust this range.

Let me know if you hit any snags—happy to help you refine this further!

内容的提问来源于stack exchange,提问作者Jara de la Court

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 12:28:15