如何基于另一DataFrame列名匹配索引筛选Pandas DataFrame行?
问题:保留df中索引与meth_clin列名匹配的行
需求
仅保留df中索引与meth_clin列名匹配的行。
用户代码
meth_clin = meth_clin.sort_index() subtype = pd.DataFrame(meth_clin["subtype"]) subtype = subtype.T subtype.columns = subtype.columns.str[:-1] i = meth_clin.iloc[:,7:].columns.str.split("|").str[0] i = [ j for j in i if "?" not in j ] a = df[df.columns.intersection(subtype.columns)] b = df[df.index.intersection(i)] b
报错信息
KeyError: "None of [Index(['ARID3A', 'ARID5B', 'ARNT', 'ARNT2', 'ATF3', 'ATOH8', 'BARX2', 'BATF', 'BATF2', 'BATF3', ... 'ZNF697', 'ZNF714', 'ZNF771', 'ZNF775', 'ZNF787', 'ZNF789', 'ZNF83', 'ZNF837', 'ZNF844', 'ZNF853'], dtype='object', length=251)] are in the [columns]"
数据说明
- df.index为包含594个元素的索引
- i为包含大量基因名称的列表
错误原因
直接用df[df.index.intersection(i)]是错误写法:pandas里方括号[]默认用来筛选列,程序会把你传入的索引集合当成列名去查找,自然找不到对应列,因此抛出KeyError。
修正代码
要筛选行,需用.loc索引器指定行的筛选条件,修正后的代码如下:
meth_clin = meth_clin.sort_index() subtype = pd.DataFrame(meth_clin["subtype"]) subtype = subtype.T subtype.columns = subtype.columns.str[:-1] i = meth_clin.iloc[:,7:].columns.str.split("|").str[0] i = [ j for j in i if "?" not in j ] a = df[df.columns.intersection(subtype.columns)] # 用.loc选取匹配索引的行 b = df.loc[df.index.intersection(i)] b
额外检查建议
可以先确认匹配的索引数量,避免筛选后得到空数据集:
common_index = df.index.intersection(i) print(f"匹配的索引数量:{len(common_index)}") b = df.loc[common_index]
内容的提问来源于stack exchange,提问作者melolilili
相关产品推荐
相关产品推荐

