在R中合并两个数据集后,基于SegmentDuration_Seconds填充Activity code求助
Got it, let's work through this together! It sounds like after merging your two time-based datasets, you need to repeat each Activity code value for exactly the number of rows specified in SegmentDuration_Seconds—just like Excel's "fill down" but tied to a count column. This is totally doable in R, and I'll walk you through two straightforward methods: one with the tidyverse (super intuitive for data manipulation) and one with base R if you prefer avoiding extra packages.
First, Let's Assume Your Merged Data Structure
To make this concrete, let's use a sample dataset that matches what you described:
# Example merged data (adjust to match your actual columns) merged_data <- data.frame( Segment_Activity = c("Morning Walk", "Midday Run", "Afternoon Rest"), SegmentDuration_Seconds = c(6, 4, 5), # Number of rows each Activity code should fill Activity_code = c("W001", "R002", "S003"), TIMEDATE = seq.POSIXt(as.POSIXct("2024-05-20 08:00:00"), by = "sec", length.out = 15) )
Method 1: Tidyverse (tidyr::uncount)
This is my go-to for this task—it's clean and explicitly does what you need. The uncount() function repeats rows based on a numeric column, and automatically carries over all other column values (including Activity_code):
library(tidyverse) # Fill Activity_code by repeating rows per SegmentDuration_Seconds filled_data <- merged_data %>% uncount(SegmentDuration_Seconds) %>% # Optional: Add a column to track position within each segment (if needed) group_by(Segment_Activity) %>% mutate(Second_In_Segment = row_number()) %>% ungroup()
After running this, each Activity_code will be repeated exactly SegmentDuration_Seconds times, just like filling down in Excel.
Method 2: Base R (No Extra Packages)
If you don't want to load the tidyverse, you can use base R's rep() function to create row indices and subset your data:
# Create repeated row indices based on SegmentDuration_Seconds row_indices <- rep(seq(nrow(merged_data)), merged_data$SegmentDuration_Seconds) # Subset the data to fill Activity_code filled_data <- merged_data[row_indices, ] # Reset row names (optional but cleaner) rownames(filled_data) <- NULL
This achieves the exact same result—each row from your merged data is repeated the number of times specified in SegmentDuration_Seconds, so Activity_code is automatically filled across those rows.
Notes for Edge Cases
- If you have
NAvalues inSegmentDuration_Seconds, addna.omit(merged_data)before running the code to exclude those rows. - If
SegmentDuration_Secondshas non-integer values (e.g., 4.7), usefloor(merged_data$SegmentDuration_Seconds)to convert to whole numbers first (since you can't have a fraction of a row).
内容的提问来源于stack exchange,提问作者EmilieTX

