You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas中extractall正则提取数据列生成但值为空的问题求助

问题分析与解决方案

你的正则表达式存在几处匹配边界的错误,导致无法正确捕获数据,以下是修正方案:

问题点说明

  • \spep后未匹配空格:原正则中\spep直接衔接primary_assembly,但实际字符串中pep与primary_assembly之间有空格,导致这部分匹配失败,整个正则无法命中。
  • description部分的空格匹配冗余:原正则中description:(?P<description>[^\[]+)\s的末尾\s会要求描述文本后必须有空格,但实际描述文本与[Source:之间的空格已被[^\[]+包含,额外的\s会导致匹配失败。
  • 末尾Source部分的正则不够精准:;*和\]*的可选匹配会引入不确定性,而实际格式是明确的;Acc:和闭合]。

修正后的正则表达式

p = re.compile(
    r'(?P<PepID>[^\s]+)\spep\s+'
    r'primary_assembly:(?P<genome_version>[^:]+):'
    r'(?P<chromosome>[^:]+):'
    r'(?P<start>[^:]+):'
    r'(?P<end>[^:]+):'
    r'(?P<strand>[^\s]+)\s+'
    r'gene:(?P<gene>[^\s]+)\s+'
    r'transcript:(?P<transcript>[^\s]+)\s+'
    r'gene_biotype:(?P<gene_biotype>[^\s]+)\s+'
    r'transcript_biotype:(?P<transcript_biotype>[^\s]+)\s+'
    r'gene_symbol:(?P<gene_symbol>[^\s]+)\s+'
    r'description:(?P<description>[^\[]+?)\s*'
    r'\[Source:(?P<source>[^;]+);'
    r'Acc:(?P<accession>[^\]]+)\]'
)

代码调整建议

由于每行只有一组匹配结果,使用str.extract()比extractall()更高效,无需额外重置索引:

df = pd.concat([
    df,
    df.fullIden.str.extract(p).fillna('')
], axis=1)

测试结果

针对你的示例数据,修正后会正确提取出以下字段:

PepIDgenome_versionchromosomestartendstrandgenetranscriptgene_biotypetranscript_biotypegene_symboldescriptionsourceaccession
ENSAPLP00000008325.2CAU_duck1.014776289936-1ENSAPLG00000008650.2ENSAPLT00000009001.2protein_codingprotein_codingSHANK3SH3 and multiple ankyrin repeat domains 3HGNC SymbolHGNC:14294

内容的提问来源于stack exchange,提问作者Alexandre P Magalhães

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 04:52:13