使用Pandas开发生物信息学程序时遭遇rsID列KeyError问题
我正在为患者报告团队开发一款生物信息学程序,用于处理患者基因检测结果:通过NCBI的rsID提取特定SNP的核苷酸数据,将其与自建参考库合并、对比,若患者核苷酸为低频率罕见位点则标记Flag。但运行脚本时,上传患者文件和群体数据后出现KeyError,提示无法找到rsID列。已知患者CSV文件存在前置表头,需跳过以获取有效数据。
onClick('Upload Patient Files First') patient_data = pd.read_csv(ask_path(),) ###patient_genotype = patient_data.loc[patient_data['rsID'] == rsID]['NCBI SNP Reference'] ##Not using onClick('Upload Population Frequency Data Next') pop_ref_data = pd.read_csv(ask_path()) #Creating a dictionary of the population reference data def pop_dict(pop_ref_data): pop_ref_dict = {} for _, row in pop_ref_data.iterrows(): variant_data ={} rsID = row['rsID'] dominant_nucleotide = row['DomNucl'] recessive_nucleotide = row['RecNucl'] dominant_freq = row['DomAllele'] recessive_freq = row['RecessiveAllele'] variant_data[dominant_nucleotide]= dominant_freq variant_data[recessive_nucleotide]= recessive_freq pop_ref_dict[rsID] = variant_data return pop_ref_dict
Traceback (most recent call last):
File "C:\Users\rcthu\PycharmProjects\WorkStuff\venv\lib\site-packages\pandas\core\indexes\base.py", line 3802, in get_loc
return self._engine.get_loc(casted_key)
File "pandas_libs\index.pyx", line 138, in pandas._libs.index.IndexEngine.get_loc
File "pandas_libs\index.pyx", line 165, in pandas._libs.index.IndexEngine.get_loc
File "pandas_libs\hashtable_class_helper.pxi", line 5745, in pandas._libs.hashtable.PyObjectHashTable.get_item
File "pandas_libs\hashtable_class_helper.pxi", line 5753, in pandas._libs.hashtable.PyObjectHashTable.get_item
KeyError: 'rsID'
The above exception was the direct cause of the following exception:
Traceback (most recent call last):
File "C:\Users\rcthu\AppData\Roaming\JetBrains\PyCharmCE2022.2\scratches\Flag Process 2.12.py", line 61, in
pop_ref_row = pop_dict(pop_ref_data)
File "C:\Users\rcthu\AppData\Roaming\JetBrains\PyCharmCE2022.2\scratches\Flag Process 2.12.py", line 41, in pop_dict
rsID = row['rsID']
File "C:\Users\rcthu\PycharmProjects\WorkStuff\venv\lib\site-packages\pandas\core\series.py", line 981, in getitem
return self._get_value(key)
File "C:\Users\rcthu\PycharmProjects\WorkStuff\venv\lib\site-packages\pandas\core\series.py", line 1089, in _get_value
loc = self.index.get_loc(label)
File "C:\Users\rcthu\PycharmProjects\WorkStuff\venv\lib\site-packages\pandas\core\indexes\base.py", line 3804, in get_loc
raise KeyError(key) from err
KeyError: 'rsID'
Process finished with exit code 1
KeyError的核心原因是读取CSV时没有跳过前置表头,导致真实列名未被识别为DataFrame的列。不管是患者文件还是群体参考文件,只要存在无效前置表头,都需要在pd.read_csv中指定参数跳过对应行数,确保真实表头被正确解析。
具体修改步骤:
- 确认前置表头行数:打开CSV文件,统计需要跳过的无效前置行数(比如文件开头有1行说明文字,就填1)。
- 修改文件读取代码:
- 患者文件读取:
patient_data = pd.read_csv(ask_path(), skiprows=N) # N为需要跳过的前置行数 - 群体参考文件读取:
pop_ref_data = pd.read_csv(ask_path(), skiprows=M) # M为群体文件的前置行数
- 患者文件读取:
- 验证列名正确性:读取后打印列名确认,避免后续出错:
print(patient_data.columns) print(pop_ref_data.columns)
额外建议:
- 如果不确定前置行数,可以用
header参数指定真实表头所在行号(比如真实表头在第2行,就填header=1),效果和skiprows一致。 - 处理CSV前手动检查文件结构,确保列名与代码中引用的
rsID、DomNucl等完全匹配(注意大小写、空格等细节)。
内容的提问来源于stack exchange,提问作者ClarkThark

