You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言合并DataFrame后出现重复条目问题求助

Fixing Duplicate Rows When Merging DataFrames in R

Hey there! Let's break down why you're seeing duplicate entries after merging your DataFrames, and how to fix it.

Why Duplicates Happen

From your code, I suspect two main culprits:

  • Non-unique dates in U.NO2.ab.03: Using unique(NO2.ab.03) only removes rows where all columns are identical. If NO2.ab.03 has multiple rows with the same Date.Local but different values in other columns, unique() will keep all those rows. When you merge, this creates a cartesian product for each matching date, leading to duplicates.
  • Implicit merge key: You didn't specify the by argument in merge(). While R will auto-match columns with the same name (here Date.Local), this can lead to unexpected behavior if there are hidden duplicates in either DataFrame's key column.

Step-by-Step Fix

  1. Ensure Unique Dates in Your NO2 Data
    First, clean NO2.ab.03 to make sure each Date.Local has only one corresponding value. Choose an aggregation method that makes sense for your data (e.g., first(), mean(), or sum()):

    # Using dplyr for cleaner aggregation (install if needed: install.packages("dplyr"))
    library(dplyr)
    
    U.NO2.ab.03 <- NO2.ab.03 %>%
      group_by(Date.Local) %>%
      summarise(across(everything(), first)) # Replace `first()` with sum/mean if needed
    

    If you don't want to use dplyr, base R works too:

    # Keep the first occurrence of each Date.Local
    U.NO2.ab.03 <- NO2.ab.03[!duplicated(NO2.ab.03$Date.Local), ]
    
  2. Merge Explicitly with Clean Data
    Now merge with a clear by parameter, and use all.x = TRUE to keep all dates from FiresNearLA.ab.03 (your original left join intent):

    # Explicitly merge on Date.Local, keep all rows from FiresNearLA.ab.03
    ind <- merge(FiresNearLA.ab.03, U.NO2.ab.03, by = "Date.Local", all.x = TRUE)
    
    # Replace NA values with 0
    ind[is.na(ind)] <- 0
    
  3. Verify No Duplicates Remain
    If you still see duplicates, check which dates are causing the issue:

    # Find duplicate dates
    dup_dates <- ind$Date.Local[duplicated(ind$Date.Local)]
    print(dup_dates)
    
    # Inspect those rows to trace the source
    ind[ind$Date.Local %in% dup_dates, ]
    

    This will help you see if duplicates are coming from FiresNearLA.ab.03 instead, and you can clean that DataFrame similarly.


内容的提问来源于stack exchange,提问作者JaElf

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 04:09:46