R语言使用lubridate处理含缺失日信息的混合格式日期及自定义输出列名问题求助
Hey there! Let's work through your R date formatting issues step by step—both the dynamic column name problem and the mixed date parsing hiccup.
1. Fixing the Dynamic Output Column Name Issue
The reason your function always creates a column named output_col is that when you write output_col = date_correction_row(input_col) inside mutate(), R treats output_col as a literal string instead of the variable name you passed to the function. To fix this, we need to use quasiquotation (a tidyverse feature that lets you pass variable names as function arguments).
Here's the revised date_correction function using the {{ }} curly-curly operator (the simplest way to handle this in modern dplyr):
date_correction <- function(df, input_col, output_col){ mutate(df, {{ output_col }} := date_correction_row({{ input_col }})) }
You can also use the older enquo() + !! syntax if you prefer (it does the same thing):
date_correction <- function(df, input_col, output_col){ input_col <- enquo(input_col) output_col <- enquo(output_col) mutate(df, !!output_col := date_correction_row(!!input_col)) }
Now when you call your function like this:
df_dates %>% date_correction(Date_original, date_formatted) %>% view()
It will generate a column named date_formatted instead of the hardcoded output_col.
2. Optimizing Mixed Date Format Parsing
Your original parse_date_time() call had a syntax error: the orders parameter expects a vector of formats, not multiple separate arguments. You wrote parse_date_time(input_col, orders="mdy", "my"), which treats "my" as an argument for tz (timezone) instead of part of the date formats.
The fix is simple—pass both formats as a single vector:
parse_date_time(input_col, orders = c("mdy", "my"))
Lubridate will automatically try each format on every date string, and fill in the missing day with the 1st of the month (which is exactly what you want).
Even Cleaner Solution (No Custom Functions Needed)
You can skip writing all those helper functions entirely and handle everything in one dplyr pipeline:
library(tidyverse) library(lubridate) # Create your sample data Observation <- seq(1:5) Date_original <- c("October 2014","August 2014","June 2013", "June 24, 2010","January 2005") df_dates <- data.frame(Observation, Date_original) # Clean dates in one step df_dates_clean <- df_dates %>% mutate( date_formatted = parse_date_time(Date_original, orders = c("mdy", "my")) %>% as.Date() # Convert from POSIXct to pure Date type (optional but cleaner) ) # Check the result df_dates_clean %>% view()
If you're working with large datasets, lubridate::parse_date_time2() is a faster alternative to parse_date_time()—it uses the same syntax but has better performance for multiple formats:
mutate( date_formatted = parse_date_time2(Date_original, orders = c("mdy", "my")) %>% as.Date() )
内容的提问来源于stack exchange,提问作者c_dinosaur

