Python:DataFrame多列数据匹配及四舍五入检测问题求助
解决方案:DataFrame文本匹配与数值四舍五入校验
一、解决SW Value包含Identifier的匹配问题
你当前的代码逻辑方向没问题,但大概率是隐藏字符(比如首尾空格、制表符、不可见Unicode字符)或格式差异(比如全角/半角符号、字符串转义残留)导致匹配失败。可以按以下方式优化:
优化后的匹配函数
def is_identifier_contained(identifier, sw_value): # 先清理首尾空白,再统一转小写 clean_id = str(identifier).strip().lower() clean_sw = str(sw_value).strip().lower() # 处理全角/半角、特殊字符统一格式(可选但实用) import unicodedata clean_id = unicodedata.normalize('NFKC', clean_id) clean_sw = unicodedata.normalize('NFKC', clean_sw) return clean_id in clean_sw
批量处理替代逐行循环(更高效)
如果数据量较大,建议用apply批量处理,避免逐行循环的低效:
import pandas as pd import unicodedata def check_identifier(row): if row["start_numerical_count"] == 1 and row["end_numerical_count"] == 0 and row["identifier_count"] == 1: clean_id = unicodedata.normalize('NFKC', str(row["Identifier"]).strip().lower()) clean_sw = unicodedata.normalize('NFKC', str(row["SW Value"]).strip().lower()) return "Yes" if clean_id in clean_sw else "No" else: return "N/A" # 假设你的DataFrame名为table table["Identifier in SW Value"] = table.apply(check_identifier, axis=1)
二、Start值四舍五入至4位小数的匹配校验
针对“匹配情况”的两种常见场景,给出实现代码:
场景1:四舍五入后与目标列值匹配
# 新增列存储四舍五入后的Start值 table["Start_Rounded_4"] = table["Start"].round(4) # 假设目标列为Target_Start,判断是否匹配 table["Start_Match"] = table.apply(lambda row: "Yes" if row["Start_Rounded_4"] == row["Target_Start"] else "No", axis=1)
场景2:判断原始值是否本身就是4位小数(四舍五入后与自身相等)
table["Start_4dp_Match"] = table.apply(lambda row: "Yes" if row["Start"].round(4) == row["Start"] else "No", axis=1)
匹配失败排查技巧
- 打印处理后的字符串原始格式,查看是否有隐藏字符:
print(f"Identifier: {repr(clean_id)}") print(f"SW Value: {repr(clean_sw)}") - 用
len()对比处理前后的字符串长度,确认是否清理掉了额外字符。
内容的提问来源于stack exchange,提问作者temp temp
相关产品推荐
相关产品推荐

