You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python:DataFrame多列数据匹配及四舍五入检测问题求助

解决方案:DataFrame文本匹配与数值四舍五入校验

一、解决SW Value包含Identifier的匹配问题

你当前的代码逻辑方向没问题,但大概率是隐藏字符(比如首尾空格、制表符、不可见Unicode字符)或格式差异(比如全角/半角符号、字符串转义残留)导致匹配失败。可以按以下方式优化:

优化后的匹配函数

def is_identifier_contained(identifier, sw_value):
    # 先清理首尾空白,再统一转小写
    clean_id = str(identifier).strip().lower()
    clean_sw = str(sw_value).strip().lower()
    # 处理全角/半角、特殊字符统一格式(可选但实用)
    import unicodedata
    clean_id = unicodedata.normalize('NFKC', clean_id)
    clean_sw = unicodedata.normalize('NFKC', clean_sw)
    return clean_id in clean_sw

批量处理替代逐行循环(更高效)

如果数据量较大,建议用apply批量处理,避免逐行循环的低效:

import pandas as pd
import unicodedata

def check_identifier(row):
    if row["start_numerical_count"] == 1 and row["end_numerical_count"] == 0 and row["identifier_count"] == 1:
        clean_id = unicodedata.normalize('NFKC', str(row["Identifier"]).strip().lower())
        clean_sw = unicodedata.normalize('NFKC', str(row["SW Value"]).strip().lower())
        return "Yes" if clean_id in clean_sw else "No"
    else:
        return "N/A"

# 假设你的DataFrame名为table
table["Identifier in SW Value"] = table.apply(check_identifier, axis=1)

二、Start值四舍五入至4位小数的匹配校验

针对“匹配情况”的两种常见场景,给出实现代码:

场景1:四舍五入后与目标列值匹配

# 新增列存储四舍五入后的Start值
table["Start_Rounded_4"] = table["Start"].round(4)
# 假设目标列为Target_Start,判断是否匹配
table["Start_Match"] = table.apply(lambda row: "Yes" if row["Start_Rounded_4"] == row["Target_Start"] else "No", axis=1)

场景2:判断原始值是否本身就是4位小数(四舍五入后与自身相等)

table["Start_4dp_Match"] = table.apply(lambda row: "Yes" if row["Start"].round(4) == row["Start"] else "No", axis=1)

匹配失败排查技巧

  • 打印处理后的字符串原始格式,查看是否有隐藏字符:
    print(f"Identifier: {repr(clean_id)}")
    print(f"SW Value: {repr(clean_sw)}")
    
  • 用len()对比处理前后的字符串长度,确认是否清理掉了额外字符。

内容的提问来源于stack exchange,提问作者temp temp

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 02:06:20