R语言:按最近日期合并含NA的行(保留较晚日期)
Hey there! Let's work through your problem—you need to fill NA values with data from nearby rows, keep the later date when merging, and avoid combining rows that are more than 15 days apart. Here's a straightforward pandas solution that hits all your requirements:
Step 1: Prep Your Data
First, make sure your Date column is in datetime format so we can calculate date differences properly:
import pandas as pd # Your sample dataframe df = pd.DataFrame({ 'Date': ['2010-01-01', '2010-01-02', '2010-01-05', '2010-01-07', '2010-01-20', '2010-01-25'], 'A': [pd.NA, 2, 3, pd.NA, 5, 6], 'B': [1, pd.NA, pd.NA, 4, pd.NA, 7] }) # Convert Date to datetime type df['Date'] = pd.to_datetime(df['Date'])
Step 2: Group Rows That Are Close Enough
We'll calculate the number of days between each row and the one before it, then create groups where rows are within 15 days of each other:
# Calculate days between consecutive dates df['day_gap'] = df['Date'].diff().dt.days # Create group IDs: start a new group if gap >15 days or it's the first row df['group_id'] = (df['day_gap'] > 15) | df['day_gap'].isna() df['group_id'] = df['group_id'].cumsum()
Step 3: Fill NAs & Keep the Latest Date
For each group, we'll use backward fill (bfill()) to plug NAs with values from later rows (which ensures we keep the later date), then grab the last row of each group (since that's the latest date):
# Fill NAs in each group, then take the final row of each group result = df.groupby('group_id').apply(lambda x: x.bfill().iloc[-1]).reset_index(drop=True) # Clean up helper columns result = result.drop(columns=['day_gap', 'group_id'])
Final Output
Run that code, and you'll get exactly the result you wanted:
Date A B 0 2010-01-02 2 1 1 2010-01-07 3 4 2 2010-01-20 5 <NA> 3 2010-01-25 6 7
Quick Breakdown of the Logic
bfill(): Backward fill is perfect here because it pulls values from later rows into earlier NAs, which aligns with your rule to keep the later date when merging.- 15-Day Limit: The
day_gap >15check ensures we don't merge rows that are too far apart—rows with gaps beyond the threshold stay as separate entries. - Grouping: By splitting the data into groups where rows are within 15 days, we only merge rows that logically belong together, avoiding unintended combinations.
内容的提问来源于stack exchange,提问作者Daniel

