基于前一行的apply函数应用:循环转apply及档案变更标记需求
Hey there! Let's solve your problem of tracking changes in user profiles and replacing loops with apply-based logic. Here's a step-by-step breakdown:
1. Understand the Requirement
We need to compare each row of a user's profile (for the same id) with the previous row, marking 1 if a column value changed and 0 if it stayed the same. We'll also replace manual loops with pandas' apply functionality, paired with grouping to handle per-user records.
2. Sample Data Setup
First, let's structure your sample data into a pandas DataFrame—this is the standard tool for tabular data tasks like this:
import pandas as pd data = { 'id': ['01', '01', '01'], 'date': ['----', '----', '----'], 'gender': ['male', 'male', 'male'], 'status': ['married', 'unmarried', 'unmarried'], 'name': ['---', '---', '---'], 'meal': ['veg', 'non-veg', 'non-veg'], 'smoker': ['yes', 'yes', 'no'] } df = pd.DataFrame(data)
3. Apply-Based Change Tracking
We'll use groupby to process each user's records separately, then apply a custom function to compare each row with its predecessor.
Efficient Vectorized Approach (Recommended)
This method avoids slow row-by-row loops by using pandas' built-in vector operations, wrapped in an apply on grouped data:
def mark_changes(group): # Initialize a DataFrame with 0s (default to no change) changes = pd.DataFrame(0, index=group.index, columns=group.columns) # Compare each row with the previous one using shift(1) change_mask = group != group.shift(1) # Mark changes as 1 where values differ changes[change_mask] = 1 # Optional: Force `id` to 0 since it's consistent per user group changes['id'] = 0 return changes # Apply the function to each user's group of records changes_df = df.groupby('id').apply(mark_changes).reset_index(drop=True) # Combine original data with change markers (add suffix to distinguish columns) final_result = pd.concat([df, changes_df.add_suffix('_changed')], axis=1)
Output
Running this code gives you a combined DataFrame showing both the original profile data and change status:
id date gender status name meal smoker id_changed date_changed gender_changed status_changed name_changed meal_changed smoker_changed 0 01 ---- male married --- veg yes 0 0 0 0 0 0 0 1 01 ---- male unmarried --- non-veg yes 0 0 0 1 0 1 0 2 01 ---- male unmarried --- non-veg no 0 0 0 0 0 0 1
Row-by-Row Apply (For Small Datasets)
If you specifically need row-level apply (though it's less efficient for large datasets), you can use this approach:
def compare_with_previous(row, group): # Get the previous row (use current row if it's the first in the group) prev_row = group.loc[row.name - 1] if row.name > group.index[0] else row # Return 1 for changed columns, 0 otherwise return (row != prev_row).astype(int) # Apply to each row in the user group changes_df = df.groupby('id').apply(lambda g: g.apply(lambda row: compare_with_previous(row, g), axis=1)).reset_index(drop=True) # Combine with original data final_result = pd.concat([df, changes_df.add_suffix('_changed')], axis=1)
4. Key Notes
- Grouping by
id: Ensures we only compare rows for the same user, not across different users. - Shift Logic:
shift(1)is pandas' built-in way to fetch the previous row in a group, eliminating the need for manual loop indexing. - Efficiency: The vectorized approach is vastly faster for large datasets compared to row-by-row loops or apply.
内容的提问来源于stack exchange,提问作者Shikhar Sharma

