如何从Pandas列中提取子字符串?获取列内每个字符串首部分
Pandas提取Cell_type列字符串的第一部分
需求说明
需要从Pandas DataFrame的Cell_type列中,提取每个字符串的第一部分内容,并对整个列批量处理,而非仅获取单个元素。
数据示例
meta.iloc[1:5] pd.DataFrame({'Assay Type': {'SRR9200814': 'RNA-Seq', 'SRR9200815': 'RNA-Seq', 'SRR9200816': 'RNA-Seq', 'SRR9200817': 'RNA-Seq'}, 'Cell_type': {'SRR9200814': 'normal neural stem cells', 'SRR9200815': 'normal neural stem cells', 'SRR9200816': 'normal neural stem cells', 'SRR9200817': 'normal neural stem cells'}, 'Type': {'SRR9200814': 'diploid', 'SRR9200815': 'diploid', 'SRR9200816': 'diploid', 'SRR9200817': 'diploid'}})
当前问题
使用代码meta["Cell_type"].str.rstrip(" ")[0]仅能获取第一个元素的完整字符串:
'normal neural stem cells'
解决方案
方法1:分割字符串取首段
利用Pandas的str.split()方法按空格分割字符串,再通过str[0]提取分割后的第一个元素,实现整列批量处理:
# 生成新列存储提取结果 meta['Cell_type_first'] = meta['Cell_type'].str.split().str[0] # 直接输出期望格式(索引命名为Run) result = meta['Cell_type'].str.split().str[0] result.index.name = 'Run' print(result)
方法2:正则提取首段
如果需要适配多空格等复杂场景,用str.extract提取开头的连续非空格字符段:
result = meta['Cell_type'].str.extract(r'^(\S+)', expand=False) result.index.name = 'Run' print(result)
期望输出
Run SRR9200813 normal SRR9200814 glioblastoma SRR9200815 normal SRR9200816 normal SRR9200817 normal
内容的提问来源于stack exchange,提问作者Anon
相关产品推荐
相关产品推荐

