如何在dplyr调用的函数中返回两个数据框?附部分函数代码
First off, let's clear up a key point about R functions: they can only send back one object at a time. So to return two data frames, we'll wrap them up in a list—that's the standard approach here. Let's fix up your existing code first, then show a more efficient dplyr-style alternative since loops can drag with large datasets.
Step 1: Finish Your Loop-Based Code
Your current code is tracking occupied dates and overwrite counts—let's complete that logic and add a second data frame (we'll use the original testdf with an extra column to track how many times each row's dates were overwritten).
library(dplyr) correct_occupancy <- function(testdf) { # Initialize first data frame: tracks detailed occupied date info occupied_dates_df <- data.frame( OccupiedDate = as.POSIXct(character()), Source = integer(), created = as.POSIXct(character()), Overwritten_cnt = integer(), testdf_row = integer(), stringsAsFactors = FALSE ) # Initialize second data frame: original data + overwrite count per row updated_testdf <- testdf %>% mutate(Overwritten_times = 0, .before = 1) # Add column to track overwrites for each row ovrd_cnt <- 0 for (i in 1:nrow(testdf)) { # Generate date range for current booking (check-in to day before check-out) testspan <- seq(as.Date(testdf$check_in[i]), as.Date(testdf$check_out[i]) - 1, by = "days") # Convert to a data frame with metadata testspan_df <- testspan %>% as.data.frame() %>% rename(OccupiedDate = ".") %>% mutate( OccupiedDate = as.POSIXct(OccupiedDate), Source = i, created = Sys.time(), Overwritten_cnt = 0, testdf_row = i ) # Check each date in the current span existing_dates <- occupied_dates_df$OccupiedDate for (j in 1:nrow(testspan_df)) { current_date <- testspan_df$OccupiedDate[j] if (current_date %in% existing_dates) { # Increment overwrite count for the existing date occupied_dates_df$Overwritten_cnt[occupied_dates_df$OccupiedDate == current_date] <- occupied_dates_df$Overwritten_cnt[occupied_dates_df$OccupiedDate == current_date] + 1 # Update the overwrite count for the current row in testdf updated_testdf$Overwritten_times[i] <- updated_testdf$Overwritten_times[i] + 1 ovrd_cnt <- ovrd_cnt + 1 } else { # Add new date to the occupied dates frame occupied_dates_df <- bind_rows(occupied_dates_df, testspan_df[j, ]) } } } # Return both data frames wrapped in a list return(list( occupied_dates = occupied_dates_df, updated_test_data = updated_testdf )) }
How to Use This Function
Once you run the function, you'll get a list—extract each data frame using $:
# Example call with your test data result <- correct_occupancy(your_test_data_frame) # Get the occupied dates details occupied_dates <- result$occupied_dates # Get the updated original data with overwrite counts updated_test_df <- result$updated_test_data
Step 2: Optimize with Dplyr/Tidyr (No Loops!)
Loops work, but they're slow for large datasets. Here's a vectorized, dplyr-friendly version that does the same job faster:
library(dplyr) library(tidyr) library(purrr) correct_occupancy_vectorized <- function(testdf) { # 1. Generate all occupied dates with row metadata all_occupied_dates <- testdf %>% mutate( # Create date range for each booking OccupiedDate = map2(check_in, check_out, ~seq(as.Date(.x), as.Date(.y)-1, by = "days")), Source = row_number(), # Track which row the date came from created = Sys.time() ) %>% unnest(OccupiedDate) %>% # Expand to one row per date mutate(OccupiedDate = as.POSIXct(OccupiedDate)) # 2. Build the occupied dates data frame with overwrite counts occupied_dates_df <- all_occupied_dates %>% group_by(OccupiedDate) %>% mutate(Overwritten_cnt = n() - 1) %>% # n()-1 because first occurrence has 0 overwrites ungroup() %>% select(OccupiedDate, Source, created, Overwritten_cnt, testdf_row = Source) # 3. Build the updated test data frame with row-level overwrite counts updated_testdf <- all_occupied_dates %>% group_by(Source) %>% # Count how many times this row's dates were duplicated (overwritten) summarize(Overwritten_times = sum(duplicated(OccupiedDate)), .groups = "drop") %>% # Join back to original testdf to retain all columns right_join(testdf %>% mutate(Source = row_number()), by = "Source") %>% relocate(Overwritten_times, .before = 1) %>% select(-Source) # Remove the temporary Source column # Return both data frames in a list list( occupied_dates = occupied_dates_df, updated_test_data = updated_testdf ) }
This version uses map2 to generate date ranges, unnest to expand into long format, and grouping to calculate counts—all operations optimized for speed in tidyverse packages.
内容的提问来源于stack exchange,提问作者Snehal Kakade

