You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在两个Pandas Series中忽略大小写查找相同条目

How to Compare Pandas Series Case-Insensitively for Matching Entries

Got it, you need to adjust your existing Pandas logic to catch cases where Identifier_1 and Identifier_2 have the same content but different capitalization—like matching 'aBc' with 'Abc'. Here's how to tweak your code to make that happen:

Modified Code

import numpy as np
import pandas as pd

# Update your condition to ignore case
df['Comment'] = np.where(
    df['Identifier_1'].str.lower() == df['Identifier_2'].str.lower(),
    'The same',
    df['Comment']
)

What's Changed?

Instead of comparing raw values directly, we convert both series to lowercase first using str.lower(). This strips out any capitalization differences, so we only check if the underlying text is identical.

If you're working with non-English characters (like German ß or Turkish dotted i), str.casefold() is a better pick—it handles Unicode case conversions more comprehensively:

df['Comment'] = np.where(
    df['Identifier_1'].str.casefold() == df['Identifier_2'].str.casefold(),
    'The same',
    df['Comment']
)

Handling Missing Values

If your series might contain NaN values, note that str.lower() leaves them as NaN, and NaN == NaN evaluates to False in Pandas. If you want to treat matching NaNs as matches too, extend the condition:

# Match case-insensitively OR both are NaN
match_condition = (
    df['Identifier_1'].str.casefold() == df['Identifier_2'].str.casefold()
) | (df['Identifier_1'].isna() & df['Identifier_2'].isna())

df['Comment'] = np.where(match_condition, 'The same', df['Comment'])

Example Test Case

Let’s test this with sample data:

# Create test DataFrame
df = pd.DataFrame({
    'Identifier_1': ['aBc', 'XyZ', 'Test', np.nan],
    'Identifier_2': ['Abc', 'xyz', 'TEST', np.nan],
    'Comment': ['', '', '', '']
})

# Apply the case-insensitive check
match_condition = df['Identifier_1'].str.casefold() == df['Identifier_2'].str.casefold()
df['Comment'] = np.where(match_condition, 'The same', df['Comment'])

print(df)

Output:

Identifier_1 Identifier_2    Comment
0          aBc          Abc  The same
1          XyZ          xyz  The same
2          Test          TEST  The same
3           NaN           NaN  The same

内容的提问来源于stack exchange,提问作者Krzysztof Pagacz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 13:12:28