基于出现频率替换Pandas行内单个值及KeyError报错解决
问题:基于特征频率生成共识表并解决KeyError报错
需求说明
- 处理特征表,规则为:针对每一列,保留在输入序列中出现次数超过半数的特征,其余替换为
X,最终删除重复行 - 输入序列共7条,因此半数阈值为
7/2=3.5,即出现次数≥4次的特征才保留
输入表格(代码中为all_res)
| Variable | 0 | 1 | 2 |
|---|---|---|---|
| Size | small | large | medium |
| Hydro | hydrophobic | hydrophobic | neutral |
| Side | sulfur | aliphatic | hydroxyl |
| Size | small | small | medium |
| Hydro | hydrophilic | hydrophobic | neutral |
| Side | basic | aliphatic | hydroxyl |
| Size | small | very large | medium |
| Hydro | hydrophobic | hydrophobic | hydrophobic |
| Side | sulfur | aromatic | aromatic |
预期输出表格
| Variable | 0 | 1 | 2 |
|---|---|---|---|
| Size | small | X | medium |
| Hydro | hydrophobic | hydrophobic | neutral |
| Side | sulfur | aliphatic | hydroxyl |
当前代码与报错
实现代码
for col in all_res: for i in all_res[col]: if all_res[col].value_counts()[i] <= len(input_sequences)/2: all_res.replace({col:i}, "X", inplace = True)
报错信息
Traceback (most recent call last): File "C:\Users\Josh\anaconda3\lib\site-packages\pandas\core\indexes\base.py", line 3621, in get_loc return self._engine.get_loc(casted_key) File "pandas\_libs\index.pyx", line 136, in pandas._libs.index.IndexEngine.get_loc File "pandas\_libs\index.pyx", line 163, in pandas._libs.index.IndexEngine.get_loc File "pandas\_libs\hashtable_class_helper.pxi", line 5198, in pandas._libs.hashtable.PyObjectHashTable.get_item File "pandas\_libs\hashtable_class_helper.pxi", line 5206, in pandas._libs.hashtable.PyObjectHashTable.get_item KeyError: 'Aliphatic' The above exception was the direct cause of the following exception: Traceback (most recent call last): File "C:\Users\Josh\OneDrive\Documents\PhD project\Thesis\Results\3. CKBP modelling project 260123 -\9. Combining data from analyses 1-6\CIWI design\ciwi.py", line 63, in <module> if all_res[col].value_counts()[i] <= len(input_sequences)/2: File "C:\Users\Josh\anaconda3\lib\site-packages\pandas\core\series.py", line 958, in __getitem__ return self._get_value(key) File "C:\Users\Josh\anaconda3\lib\site-packages\pandas\core\series.py", line 1069, in _get_value loc = self.index.get_loc(label) File "C:\Users\Josh\anaconda3\lib\site-packages\pandas\core\indexes\base.py", line 3623, in get_loc raise KeyError(key) from err KeyError: 'Aliphatic'
报错原因
- 循环中修改原数据:使用
inplace=True直接修改原表,导致后续循环时,原有特征值已被替换为X,再去value_counts()中查找原数值会出现KeyError - 大小写不匹配:报错中的
'Aliphatic'与表格中的'aliphatic'大小写不一致,触发索引查找失败
解决方案
import pandas as pd # 计算阈值:出现次数超过半数才保留 threshold = len(input_sequences) / 2 # 复制原表,避免修改原数据引发的问题 processed_df = all_res.copy() for col in processed_df.columns: # 提前计算当前列所有值的出现次数 count_series = processed_df[col].value_counts() # 筛选出需要替换为X的特征值 values_to_replace = count_series[count_series <= threshold].index.tolist() # 批量替换符合条件的值 processed_df[col] = processed_df[col].apply(lambda x: "X" if x in values_to_replace else x) # 删除重复行,得到最终结果 final_df = processed_df.drop_duplicates() print(final_df)
代码说明
- 先复制原表,避免循环中修改原数据导致的索引异常
- 提前计算每列的特征出现次数,一次性筛选出所有需要替换的值,批量替换效率更高且不会触发KeyError
- 最后用
drop_duplicates()删除重复行,得到目标共识表
内容的提问来源于stack exchange,提问作者Quockerwodger
相关产品推荐
相关产品推荐

