用Pandas Apply替代双重循环,将文本分析分数整合至DataFrame
Let's break down how to solve this problem properly—ditching the slow nested loops and fixing that TypeError you ran into. First, your core goal is to map each CODE in your main df to its corresponding phrase scores from the scores dataframe, then pivot those scores into wide-format columns (like Phrase_stove) and populate them efficiently.
Optimal Solution: Vectorized Operations (No Apply/Loops)
The fastest way to do this in pandas is using pivot + merge, which leverages pandas' built-in vectorized operations (way faster than any loop or apply approach, especially with large datasets). Here's how:
import pandas as pd # Your original data scores = pd.DataFrame(columns=['code','phrase','score'], data=[['01A','stove',0.673], ['01A','hot',0.401], ['XR3','service',0.437], ['XR3','stove',0.408], ['0132','replace',0.655], ['0132','recommend',0.472]]) df = pd.DataFrame(columns=['CODE','YR_OPEN','COST'], data=[['01A',2004,173.23],['01A',2008,82.18], ['01A',2012,939.32],['01A',2010,213.21], ['01A',2016,173.39],['01A',2013,183.46], ['XR3',2017,998.61],['XR3',2012,38.99], ['XR3',2017,923.71],['XR3',2004,832.23], ['0132',2004,823.12],['0132',2017,832.12], ['0132',2002,887.51],['0132',2002,92.35], ['0132',2013,21.03],['0132',2008,9472.94], ['0132',2012,341.93],['0132',2008,881.36]]) # Step 1: Pivot scores into a wide-format table (code as index, phrases as columns) scores_wide = scores.pivot( index='code', columns='phrase', values='score' ).fillna(0) # Fill missing phrases with 0 scores_wide = scores_wide.add_prefix('Phrase_') # Rename columns to match your desired format # Step 2: Merge the wide scores table with your original df result = df.merge(scores_wide, left_on='CODE', right_index=True, how='left') # Ensure any remaining gaps are filled with 0 (shouldn't be needed, but safe) result = result.fillna(0) print(result)
This produces exactly the output you want, and it’s orders of magnitude faster than loops because it uses pandas' optimized backend instead of iterating row-by-row.
Fixing Your Apply Attempts & TypeError
If you absolutely need to use apply (though it’s not recommended for performance), let’s fix your code and explain the error:
The TypeError Explained
Your error apply() got multiple values for argument 'axis' happened because you passed row['rowIndex'] as a positional argument to apply(), but pandas interprets positional arguments after the function as part of the args parameter. You need to wrap extra arguments in a tuple using the args keyword parameter.
Improved Apply Approach (Still Not As Fast As Vectorized)
First, initialize all phrase columns with 0, then use apply to populate the scores for each row:
# Initialize all phrase columns with 0 unique_phrases = scores['phrase'].unique() for phrase in unique_phrases: df[f'Phrase_{phrase}'] = 0 # Define a function to return scores for a given CODE def populate_scores(row): # Get scores for this row's CODE code_scores = scores[scores['code'] == row['CODE']] # Create a series to hold scores for all phrases score_updates = pd.Series(0, index=[f'Phrase_{p}' for p in unique_phrases]) for _, score_row in code_scores.iterrows(): score_updates[f'Phrase_{score_row["phrase"]}'] = score_row['score'] return score_updates # Apply the function and merge back to the original df score_updates_df = df.apply(populate_scores, axis=1) result = pd.concat([df, score_updates_df], axis=1)
Note: This avoids modifying the original df inside the apply function (which is a common source of bugs) and instead returns a series of updates to concatenate.
Why Your Original Apply Failed
In your attempt_helper call, you should have passed the row index via args:
# Fixed version of your helper call (not recommended, but just for explanation) score_subset.apply(attempt_helper, args=(row['rowIndex'],), axis=1)
But even with this fix, modifying df directly inside apply is risky because pandas may process rows out of order, leading to incorrect values.
Final Recommendation
Stick with the pivot + merge approach—it’s clean, fast, and idiomatic pandas. Loops and apply should be your last resort for row-wise operations, as they don’t take advantage of pandas' optimized performance.
内容的提问来源于stack exchange,提问作者3pitt

