You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python DataFrame提取冒号与连字符间字符串报错排查

问题原因

re.findall() 只能处理单个字符串/字节对象,但你直接把整个Pandas Series传了进去——哪怕用astype(str)转了类型,返回的还是Series,不是单个字符串,这就触发了类型错误。

两种解决办法

方法1:用Pandas原生的str.extract()(最推荐)

Pandas专门提供了字符串处理方法,可以批量处理Series里的每一个元素,正则里的捕获组会直接提取目标内容:

all_cancers['提取的数值'] = all_cancers.iloc[:,3].astype(str).str.extract(r'\:(.*?)\-')

这里用原始字符串r''包裹正则,避免转义符出问题,(.*?)会精准捕获冒号和连字符之间的内容。

方法2:用apply()配合re.findall()

如果非要用re.findall(),就得给每个字符串单独处理,通过apply()遍历Series的每个元素:

import re

def get_target_num(s):
    matches = re.findall(r'\:(.*?)\-', s)
    return matches[0] if matches else None

all_cancers['提取的数值'] = all_cancers.iloc[:,3].astype(str).apply(get_target_num)

验证结果

处理完后查看新列:

print(all_cancers['提取的数值'])

会得到:

0    100414771
1     10506157
2    109655506
3    113903257
4    117598869
Name: 提取的数值, dtype: object

内容的提问来源于stack exchange,提问作者Achal Neupane

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 23:35:19