You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何Python中外观相同的非英文字符串不相等?Pandas列名匹配解决方法

解决Pandas非英文字符列名匹配问题

问题场景

使用Pandas的pd.read_csv读取CSV文件时,手动输入的列名"Mã NPP"和DataFrame中的列名外观完全一致,但字符串比对返回False,导致无法通过列名筛选数据。

代码示例

import pandas as pd  
df = pd.read_csv(r'https://raw.githubusercontent.com/leanhdung1994/data/main/DIST_INVOICE_RETURN.csv', encoding = 'utf8', header = 0, nrows = 1)
x = 'Mã NPP'
y = df.columns[0]
print(x, '\n', y)
print(x == y)

运行结果

Mã NPP 
 Mã NPP
False

核心原因

这种现象是因为非英文字符存在不同的Unicode表示形式:比如ã可能是单个合成字符(Unicode码点U+00E3),也可能是基础字符a(U+0061)加上组合重音符̃(U+0303)拼接而成。两种形式外观完全相同,但底层编码不同,导致字符串比对不相等。

解决方法

1. 排查字符编码差异

通过遍历字符的Unicode码点,确认差异所在:

# 打印每个字符的Unicode码点
print([ord(c) for c in x])
print([ord(c) for c in y])

运行后会看到两个字符串的码点列表不同,比如一个是[77, 227, 32, 78, 80, 80],另一个是[77, 97, 771, 32, 78, 80, 80],其中771就是组合重音符的码点。

2. 标准化Unicode字符串

使用Python标准库unicodedata的normalize方法,将字符串统一为**NFC(标准等价合成)**格式,把组合字符合并为单个合成字符:

import unicodedata

# 标准化手动输入的字符串和DataFrame列名
x_normalized = unicodedata.normalize('NFC', x)
y_normalized = unicodedata.normalize('NFC', y)
print(x_normalized == y_normalized)  # 输出True

3. 批量标准化DataFrame列名

读取CSV后直接对所有列名做标准化处理,后续即可用手动输入的列名正常筛选:

import pandas as pd
import unicodedata

df = pd.read_csv(r'https://raw.githubusercontent.com/leanhdung1994/data/main/DIST_INVOICE_RETURN.csv', encoding='utf8', header=0, nrows=1)
# 对所有列名执行Unicode标准化
df.columns = [unicodedata.normalize('NFC', col) for col in df.columns]

# 现在可以正常筛选列
target_col = 'Mã NPP'
print(df[target_col])

4. 直接从DataFrame获取列名(避免手动输入)

如果不需要手动输入列名,直接从DataFrame的columns属性中获取目标列名,完全规避编码差异问题:

target_col = df.columns[0]
print(df[target_col])

内容的提问来源于stack exchange,提问作者Akira

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 09:45:40