按第2个下划线分割DataFrame索引触发TypeError的问题求助
解决Pandas索引按第二个下划线分割的TypeError问题
需求
将anno DataFrame的索引按第2个下划线(_)分隔符分割,保留分割后的前两部分并重新赋值给索引。
错误代码及报错
用户编写的错误代码:
anno.index = ','.join(str(v) for v in anno.index) '_'.join(anno.index.split('_')[:2])
触发的TypeError回溯:
--------------------------------------------------------------------------- TypeError Traceback (most recent call last) Input In [66], in <cell line: 2>() 1 anno = pd.read_csv(directory_path + "GSE159115_ccRCC_anno.csv.gz", compression="gzip", header=0, sep=",", on_bad_lines="warn", index_col=0) ----> 2 anno.index = ','.join(str(v) for v in anno.index) 3 '_'.join(anno.index.split('_')[:2]) File ~/.local/lib/python3.9/site-packages/pandas/core/generic.py:6002, in NDFrame.__setattr__(self, name, value) 6000 try: 6001 object.__getattribute__(self, name) -> 6002 return object.__setattr__(self, name, value) 6003 except AttributeError: 6004 pass File ~/.local/lib/python3.9/site-packages/pandas/_libs/properties.pyx:69, in pandas._libs.properties.AxisProperty.__set__() File ~/.local/lib/python3.9/site-packages/pandas/core/generic.py:729, in NDFrame._set_axis(self, axis, labels) 723 @final 724 def _set_axis(self, axis: AxisInt, labels: AnyArrayLike | list) -> None: 725 """ 726 This is called from the cython code when we set the `index` attribute 727 directly, e.g. `series.index = [1, 2, 3]`. 728 """ -> 729 labels = ensure_index(labels) 730 self._mgr.set_axis(axis, labels) 731 self._clear_item_cache() File ~/.local/lib/python3.9/site-packages/pandas/core/indexes/base.py:7128, in ensure_index(index_like, copy) 7126 return Index(index_like, copy=copy, tupleize_cols=False) 7127 else: -> 7128 return Index(index_like, copy=copy) File ~/.local/lib/python3.9/site-packages/pandas/core/indexes/base.py:516, in Index.__new__(cls, data, dtype, copy, name, tupleize_cols) 513 data = com.asarray_tuplesafe(data, dtype=_dtype_obj) 515 elif is_scalar(data): -> 516 raise cls._raise_scalar_data_error(data) 517 elif hasattr(data, "__array__"): 518 return Index(np.asarray(data), dtype=dtype, copy=copy, name=name) File ~/.local/lib/python3.9/site-packages/pandas/core/indexes/base.py:5066, in Index._raise_scalar_data_error(cls, data) 5061 @final 5062 @classmethod 5063 def _raise_scalar_data_error(cls, data): 5064 # We return the TypeError so that we can raise it from the constructor 5065 # in order to keep mypy happy -> 5066 raise TypeError( 5067 f"{cls.__name__}(...) must be called with a collection of some " 5068 f"kind, {repr(data)} was passed" 5069 ) TypeError: Index(...) must be called with a collection of some kind, 'SI_18854_AAACCTGCAAGTAGTA-1,SI_18854_AAACCTGTCCACTGGG-1,SI_18854_AAACCTGTCCTTTCTC-1,SI_18854_AAACGGGCAAACTGCT-1,SI_18854_AAACGGGCAAGGTTTC-1' was passed
错误原因
- 第一行代码将整个索引转换成了单一字符串(用逗号连接所有索引值),但Pandas要求索引必须是集合类型(如列表、数组),因此触发TypeError。
- 第二行代码尝试对整个索引字符串调用
split,没有实现对每个索引元素单独分割的逻辑,完全不符合需求。
正确解决方案
方法一:使用Pandas str访问器(简洁高效)
利用Pandas内置的字符串处理方法,批量处理每个索引元素:
# 对每个索引元素按下划线分割最多2次,取前两部分再拼接 anno.index = anno.index.str.split('_', n=2).str[:2].str.join('_')
str.split('_', n=2):将每个索引字符串按下划线分割最多2次,得到包含3个元素的列表(例如['SI', '18854', 'AAACCTGCAAGTAGTA-1'])str[:2]:提取列表的前两个元素str.join('_'):将前两个元素用下划线重新拼接
方法二:列表推导式(直观易懂)
通过遍历每个索引元素,手动处理后生成新索引列表:
anno.index = ['_'.join(idx.split('_')[:2]) for idx in anno.index]
- 遍历原索引的每个元素
idx - 对
idx按下划线分割后取前两部分,用下划线拼接 - 生成的列表直接赋值给
anno.index
处理效果示例
原索引示例:
Index(['SI_18854_AAACCTGCAAGTAGTA-1', 'SI_18854_AAACCTGTCCACTGGG-1', 'SI_18854_AAACCTGTCCTTTCTC-1', 'SI_18854_AAACGGGCAAACTGCT-1', 'SI_18854_AAACGGGCAAGGTTTC-1'], dtype='object', name='cell', length=20748)
处理后索引:
Index(['SI_18854', 'SI_18854', 'SI_18854', 'SI_18854', 'SI_18854'], dtype='object', name='cell', length=20748)
内容的提问来源于stack exchange,提问作者Anon
相关产品推荐
相关产品推荐

