如何在R中为data.frame各列提取首值>1的指定行至新表
Solution using data.table
Here's a clean implementation to achieve your goal:
library(data.table) # Convert your data.frame to data.table (if not already done) setDT(dt) # Add a column for the actual row numbers (since your rows start at 52) dt[, row_id := as.integer(rownames(dt))] # Function to get the row range for a single column get_target_rows <- function(col_name) { # Find the first row number where the column value is >1 first_row <- dt[get(col_name) > 1, min(row_id)] # Return the range if such a row exists, else NULL if (!is.na(first_row)) { return(seq(from = first_row - 1, to = first_row + 25)) } else { return(NULL) } } # Apply the function to all columns (excluding the new row_id column) all_target_rows <- unique(unlist(lapply(names(dt)[names(dt) != "row_id"], get_target_rows))) # Subset the original data to these rows, removing the row_id column result_dt <- dt[row_id %in% all_target_rows, .SD, .SDcols = names(dt)[names(dt) != "row_id"]]
Key Step Explanations:
- Add row_id column: Captures your actual row numbers (52, 53, ...) instead of the default 1-based data frame indices, which is critical for accurate range calculation.
- get_target_rows function: For each column, finds the first row where values exceed 1, then generates the required range from
n-1ton+25. - Combine ranges: Flattens all column-specific ranges and keeps only unique rows to avoid duplicates.
- Subset data: Extracts the relevant rows from the original data, excluding the helper
row_idcolumn.
Why Your Previous Attempts Failed:
- Your first code used
i>1which checks if the column index is greater than 1, not if values in the column are greater than 1. You should reference column values withdf[[i]] > 1instead. - Your second code used
min(which(df[i]>1))which returns the 1-based index of the row (e.g., 3 for row 54) instead of the actual row number (54), leading to incorrect range calculation.
内容的提问来源于stack exchange,提问作者AlbertQuackstein
相关产品推荐
相关产品推荐

