如何校验pd.util.hash_pandas_object生成的各列哈希值一致
问题场景
开发以DataFrame为输入的应用时,需要校验间隔选取的奇数列内容完全一致,使用的示例DataFrame如下:
import pandas as pd df = pd.DataFrame({'store': ['Blank_A09', 'Control_4p','13_MEG3','04_GRB10','02_PLAGL1','Control_21q','01_PLAGL1','11_KCNQ10T1','16_SNRPN','09_H19','Control_6p','06_MEST'], 'quarter': [1, 1, 2, 2, 1, 1, 2, 2,2,2,2,2], 'employee': ['Blank_A09', 'Control_4p','13_MEG3','04_GRB10','02_PLAGL1','Control_21q','01_PLAGL1','11_KCNQ10T1','16_SNRPN','09_H19','Control_6p','06_MEST'], 'foo': [1, 1, 2, 2, 1, 1, 9, 2,2,4,2,2], 'columnX': ['Blank_A09', 'Control_4p','13_MEG3','04_GRB10','02_PLAGL1','Control_21q','01_PLAGL1','11_KCNQ10T1','16_SNRPN','09_H19','Control_6p','06_MEST']}) print(df)
输出的DataFrame内容:
store quarter employee foo columnX 0 Blank_A09 1 Blank_A09 1 Blank_A09 1 Control_4p 1 Control_4p 1 Control_4p 2 13_MEG3 2 13_MEG3 2 13_MEG3 3 04_GRB10 2 04_GRB10 2 04_GRB10 4 02_PLAGL1 1 02_PLAGL1 1 02_PLAGL1 5 Control_21q 1 Control_21q 1 Control_21q 6 01_PLAGL1 2 01_PLAGL1 9 01_PLAGL1 7 11_KCNQ10T1 2 11_KCNQ10T1 2 11_KCNQ10T1 8 16_SNRPN 2 16_SNRPN 2 16_SNRPN 9 09_H19 2 09_H19 4 09_H19 10 Control_6p 2 Control_6p 2 Control_6p 11 06_MEST 2 06_MEST 2 06_MEST
已编写的奇数列筛选、哈希计算代码如下:
# 筛选奇数列 df_odd = df.iloc[:,::2] # 计算各列哈希值 pd.util.hash_pandas_object(df.T, index=False)
运行后得到各列哈希结果:
store 18266754969677227875 employee 18266754969677227875 columnX 18266754969677227875 dtype: uint64
需要实现逻辑判断这些哈希值是否全部相同,从而确认目标列内容一致。
实现方案
首先修正原代码的逻辑问题:哈希计算时传入全量DataFrame的转置df.T会把不需要校验的偶数列也纳入计算,应该传入提前筛选好的奇数列转置df_odd.T。
判断哈希值全部一致有两种简单可落地的写法:
- 方法1:判断所有哈希值和第一个哈希值完全相等,用
.all()做全量匹配
hash_res = pd.util.hash_pandas_object(df_odd.T, index=False) # 返回True即所有列内容一致,False则存在不一致 is_all_same = (hash_res == hash_res.iloc[0]).all()
- 方法2:统计哈希值的唯一值数量,若唯一值数量为1则说明全部一致
hash_res = pd.util.hash_pandas_object(df_odd.T, index=False) is_all_same = hash_res.nunique() == 1
边界场景处理
如果筛选后奇数列数量不足2列(空列/仅1列),不需要做一致性校验,可以提前加判断避免索引报错:
if len(df_odd.columns) < 2: is_all_same = True else: hash_res = pd.util.hash_pandas_object(df_odd.T, index=False) is_all_same = (hash_res == hash_res.iloc[0]).all()
完整可运行测试代码:
import pandas as pd df = pd.DataFrame({'store': ['Blank_A09', 'Control_4p','13_MEG3','04_GRB10','02_PLAGL1','Control_21q','01_PLAGL1','11_KCNQ10T1','16_SNRPN','09_H19','Control_6p','06_MEST'], 'quarter': [1, 1, 2, 2, 1, 1, 2, 2,2,2,2,2], 'employee': ['Blank_A09', 'Control_4p','13_MEG3','04_GRB10','02_PLAGL1','Control_21q','01_PLAGL1','11_KCNQ10T1','16_SNRPN','09_H19','Control_6p','06_MEST'], 'foo': [1, 1, 2, 2, 1, 1, 9, 2,2,4,2,2], 'columnX': ['Blank_A09', 'Control_4p','13_MEG3','04_GRB10','02_PLAGL1','Control_21q','01_PLAGL1','11_KCNQ10T1','16_SNRPN','09_H19','Control_6p','06_MEST']}) df_odd = df.iloc[:,::2] if len(df_odd.columns) < 2: is_all_match = True else: hash_res = pd.util.hash_pandas_object(df_odd.T, index=False) is_all_match = (hash_res == hash_res.iloc[0]).all() print(is_all_match) # 示例数据下输出True
内容的提问来源于stack exchange,提问作者Manolo Dominguez Becerra
相关产品推荐
相关产品推荐

