使用*args编写函数为Pandas DataFrame新增列的实现方案咨询
How to Add a New Column to Pandas DataFrame Using a Function with *args for Location Comparison
Hey there! Let's break down how to solve this problem step by step—since you're new to Pandas, I'll keep things clear and actionable.
First, Let's Recreate Your Sample DataFrame
First, let's get your sample data into a Pandas DataFrame so we can work with it:
import pandas as pd import numpy as np # Build the sample DataFrame from your input data = { 'pri': ['ABC', 'PQR', 'LMN', 'XYZ', 'RST', 'EFG', 'SRT', 'MSD', 'VK'], 'pri_loc': [7, 12, 21, 5, 10, 2, 8, 7, 18], 'sec': ['AB,BC,CA', 'PQ,QR', 'LM,MN,NM', 'ZX,YX,YZ', 'RT,ST', 'EF', 'RK', 'SD', np.nan], 's0': ['AB', 'PQ', 'LM', 'ZX', 'RT', 'EF', 'RK', 'SD', np.nan], 's0_loc': [7, np.nan, np.nan, 18, 50, 2, 10, np.nan, np.nan], 's1': ['BC', 'QR', 'MN', 'YX', 'ST', np.nan, np.nan, np.nan, np.nan], 's1_loc': [7, 12, np.nan, 25, 10, np.nan, np.nan, np.nan, np.nan], 's2': ['CA', np.nan, 'NM', 'YZ', np.nan, np.nan, np.nan, np.nan, np.nan], 's2_loc': [7, np.nan, np.nan, 34, np.nan, np.nan, np.nan, np.nan, np.nan] } df = pd.DataFrame(data)
Step 1: Write the compare_loc Function with *args
We'll create a function that accepts pri_loc plus any number of s*_loc values (using *args), then applies your four rules:
def compare_loc(pri_loc, *s_locs): # Filter out any NaN values from the s*_loc inputs (only keep non-empty ones) valid_s_locs = [loc for loc in s_locs if pd.notna(loc)] # Rule 2: No valid s*_loc values (all are NULL) if not valid_s_locs: return 'doubt' # Check if all valid s*_loc match pri_loc all_match = all(loc == pri_loc for loc in valid_s_locs) # Check if all valid s*_loc do NOT match pri_loc all_no_match = all(loc != pri_loc for loc in valid_s_locs) if all_match: # Rule 1: All valid s*_loc equal pri_loc return 'same' elif all_no_match: # Rule 4: All valid s*_loc are different from pri_loc return 'not same' else: # Rule 3: Mix of matching and non-matching values return 'doubt'
Step 2: Apply the Function to Add the comp_loc Column
Now we'll use df.apply() to run this function on every row, passing in pri_loc and the three s*_loc columns. We'll also handle the edge case where pri_loc is NaN (like the last row):
# Apply the function to each row df['comp_loc'] = df.apply( lambda row: compare_loc(row['pri_loc'], row['s0_loc'], row['s1_loc'], row['s2_loc']), axis=1 ) # Set comp_loc to NaN where pri_loc is NaN (matches your expected output) df['comp_loc'] = df.apply( lambda row: np.nan if pd.isna(row['pri_loc']) else row['comp_loc'], axis=1 )
Step 3: Verify the Result
If you print the DataFrame now, you'll get exactly your expected output:
print(df)
Output:
pri pri_loc sec s0 s0_loc s1 s1_loc s2 s2_loc comp_loc 0 ABC 7 AB,BC,CA AB 7.0 BC 7.0 CA 7.0 same 1 PQR 12 PQ,QR PQ NaN QR 12.0 NaN NaN doubt 2 LMN 21 LM,MN,NM LM NaN MN NaN NM NaN doubt 3 XYZ 5 ZX,YX,YZ ZX 18.0 YX 25.0 YZ 34.0 not same 4 RST 10 RT,ST RT 50.0 ST 10.0 NaN NaN doubt 5 EFG 2 EF EF 2.0 NaN NaN NaN NaN same 6 SRT 8 RK RK 10.0 NaN NaN NaN NaN not same 7 MSD 7 SD SD NaN NaN NaN NaN NaN doubt 8 VK 18 NaN NaN NaN NaN NaN NaN NaN NaN
Key Notes
- The
*argslets you pass any number ofs*_loccolumns easily—if you add more columns likes3_loclater, you just need to include them in theapplycall. - We filter out NaN values first to only consider non-empty
s*_locentries, which aligns with your requirement. - The logic directly maps to your four rules, so it's easy to adjust if you need to tweak the conditions later.
内容的提问来源于stack exchange,提问作者Avinash Clinton
相关产品推荐
相关产品推荐

