基于多列条件连接/合并两个R语言数据框技术问询
Alright, let's work through this problem step by step. First, I spot a critical detail: your maindata has dates in 2018, while the weather dataset uses 2017. A direct join on district and date would return no matches because of that year gap. So we need to adjust the dates to align by month and day first, then perform the join.
Step 1: Start with Your Original Data
First, let's confirm the data frames you're working with (I'll keep your exact code here):
# Main dataset maindata <- data.frame( eventid = c(1:10), district = c(rep("lucknow",2), rep("allahabad",1), rep("kanpur", 2)), date = c(rep("2018-01-01", 2), rep("2018-01-02", 1), rep("2018-01-03", 2)) ) # Weather dataset weather <- data.frame( district = c(rep("lucknow", 4), rep("allahabad", 3), rep("kanpur", 3)), date = c(rep("2017-01-01", 4), rep("2017-01-02", 3), rep("2017-01-03", 3)), temperature = c(rep("19.3",2), rep("22.1",1), rep("24.1", 2)) )
Note: R automatically cycles shorter columns to match the longest one, so maindata will have 10 rows (matching eventid), and weather will have 10 rows too.
Step 2: Load Tools for Date Handling & Joins
We'll use dplyr for easy data manipulation and joins, plus lubridate to simplify date operations:
library(dplyr) library(lubridate)
Step 3: Align Dates by Month & Day
We'll create a new column in both data frames that isolates the month and day (ignoring the year). This lets us match rows where the district is the same and the date falls on the same month-day combination:
# Process maindata: convert date to Date type and extract month-day maindata <- maindata %>% mutate( date = ymd(date), # Convert character string to Date object month_day = format(date, "%m-%d") # Create "MM-DD" string ) # Process weather the same way weather <- weather %>% mutate( date = ymd(date), month_day = format(date, "%m-%d") )
Step 4: Perform the Join
Now we can join the data frames using district and month_day as our matching keys. Choose the join type based on your needs:
Option 1: Left Join (Keep all rows from maindata)
Use this if you want to retain every row in maindata, even if there's no corresponding weather data (those rows will have NA for temperature):
joined_data <- maindata %>% left_join( weather %>% select(district, month_day, temperature), # Only keep columns we need for joining by = c("district", "month_day") )
Option 2: Inner Join (Only keep matching rows)
Use this if you only want rows where there's a match in both maindata and weather:
joined_data <- maindata %>% inner_join( weather %>% select(district, month_day, temperature), by = c("district", "month_day") )
Step 5: Clean Up (Optional)
If you don't need the month_day column anymore, you can drop it with:
joined_data <- joined_data %>% select(-month_day)
Example Result
After the left join, your output will look something like this (first 5 rows):
eventid district date temperature 1 1 lucknow 2018-01-01 19.3 2 2 lucknow 2018-01-01 19.3 3 3 allahabad 2018-01-02 22.1 4 4 kanpur 2018-01-03 24.1 5 5 kanpur 2018-01-03 24.1
内容的提问来源于stack exchange,提问作者Adarsha Murthy

