如何在两个Pandas Series中忽略大小写查找相同条目
Got it, you need to adjust your existing Pandas logic to catch cases where Identifier_1 and Identifier_2 have the same content but different capitalization—like matching 'aBc' with 'Abc'. Here's how to tweak your code to make that happen:
Modified Code
import numpy as np import pandas as pd # Update your condition to ignore case df['Comment'] = np.where( df['Identifier_1'].str.lower() == df['Identifier_2'].str.lower(), 'The same', df['Comment'] )
What's Changed?
Instead of comparing raw values directly, we convert both series to lowercase first using str.lower(). This strips out any capitalization differences, so we only check if the underlying text is identical.
If you're working with non-English characters (like German ß or Turkish dotted i), str.casefold() is a better pick—it handles Unicode case conversions more comprehensively:
df['Comment'] = np.where( df['Identifier_1'].str.casefold() == df['Identifier_2'].str.casefold(), 'The same', df['Comment'] )
Handling Missing Values
If your series might contain NaN values, note that str.lower() leaves them as NaN, and NaN == NaN evaluates to False in Pandas. If you want to treat matching NaNs as matches too, extend the condition:
# Match case-insensitively OR both are NaN match_condition = ( df['Identifier_1'].str.casefold() == df['Identifier_2'].str.casefold() ) | (df['Identifier_1'].isna() & df['Identifier_2'].isna()) df['Comment'] = np.where(match_condition, 'The same', df['Comment'])
Example Test Case
Let’s test this with sample data:
# Create test DataFrame df = pd.DataFrame({ 'Identifier_1': ['aBc', 'XyZ', 'Test', np.nan], 'Identifier_2': ['Abc', 'xyz', 'TEST', np.nan], 'Comment': ['', '', '', ''] }) # Apply the case-insensitive check match_condition = df['Identifier_1'].str.casefold() == df['Identifier_2'].str.casefold() df['Comment'] = np.where(match_condition, 'The same', df['Comment']) print(df)
Output:
Identifier_1 Identifier_2 Comment 0 aBc Abc The same 1 XyZ xyz The same 2 Test TEST The same 3 NaN NaN The same
内容的提问来源于stack exchange,提问作者Krzysztof Pagacz

