R语言自定义函数开发请求:提取数据框中各日期列的最小与最大值并生成新数据框
Solution: Extract Date Column Ranges into a Summary Data Frame
Hey there! As someone who's navigated R's learning curve, I totally get how tricky it can be to dynamically wrangle date columns. Let's break this down into simple, actionable steps to build your custom function.
Implementation Approach
Here's the core logic we'll follow:
- Identify Date Columns: Automatically detect all columns in your data frame that are of
Datetype (works no matter how many date columns you have). - Calculate Min/Max for Each Date Column: For every detected date column, compute its minimum and maximum values.
- Reshape the Output: Convert the results into the wide-format data frame you specified, with clear column names like
min_Date1,max_Date1.
Complete Code Example
First, let's set up your sample data frame (replace this with your actual data):
# Load required libraries library(dplyr) library(tidyr) library(purrr) # Create sample data frame df <- tibble::tibble( Num1 = c(1, 2, 3), Chr1 = c("Name1", "Name2", "Name3"), Date1 = as.Date(c("2012-01-01", "2013-01-01", "2014-01-01")), Date2 = as.Date(c("2018-02-01", "2016-02-01", "2014-02-01")), Date3 = as.Date(c("2015-05-01", "2014-05-01", "2013-05-01")) )
Now, let's build the custom function:
get_date_summary <- function(input_df) { # Step 1: Filter only Date-type columns date_columns <- input_df %>% select(where(is.Date)) # Step 2: Compute min/max for each date column, format into long data date_stats <- date_columns %>% map_dfr( ~ tibble(min_val = min(., na.rm = TRUE), max_val = max(., na.rm = TRUE)), .id = "col_name" ) %>% pivot_longer( cols = c(min_val, max_val), names_to = "stat_type", values_to = "date_value" ) %>% mutate( # Create column names like "min_Date1", "max_Date2" summary_col = paste(gsub("_val", "", stat_type), col_name, sep = "_") ) # Step 3: Reshape to wide format and add the "date" row identifier summary_df <- date_stats %>% select(summary_col, date_value) %>% pivot_wider(names_from = summary_col, values_from = date_value) %>% mutate(row_label = "date") # Optional: Set row label as rowname (matches your expected output) rownames(summary_df) <- summary_df$row_label summary_df <- summary_df %>% select(-row_label) return(summary_df) } # Run the function on your data frame df1 <- get_date_summary(df) print(df1)
Key Details Explained
select(where(is.Date)): This uses dplyr'swhere()helper to automatically pick out all columns that are of classDate—perfect for handling variable numbers of date columns.map_dfr(...): The purrr package'smap_dfr()lets us iterate over each date column, compute min/max, and bind the results into a single data frame. The.idargument keeps track of which column each statistic comes from.pivot_longer()+pivot_wider(): These tidyr functions reshape our data from long to wide format, creating the clean column names you want.na.rm = TRUE: This ensures the function works even if your date columns have missing values (it skips NA when calculating min/max).
Expected Output
When you run the code, you'll get exactly the data frame you requested:
min_Date1 max_Date1 min_Date2 max_Date2 min_Date3 max_Date3 date 2012-01-01 2014-01-01 2014-02-01 2018-02-01 2013-05-01 2015-05-01
内容的提问来源于stack exchange,提问作者Zuzana
相关产品推荐
相关产品推荐

