Python DataFrame提取冒号与连字符间字符串报错排查
问题原因
re.findall() 只能处理单个字符串/字节对象,但你直接把整个Pandas Series传了进去——哪怕用astype(str)转了类型,返回的还是Series,不是单个字符串,这就触发了类型错误。
两种解决办法
方法1:用Pandas原生的str.extract()(最推荐)
Pandas专门提供了字符串处理方法,可以批量处理Series里的每一个元素,正则里的捕获组会直接提取目标内容:
all_cancers['提取的数值'] = all_cancers.iloc[:,3].astype(str).str.extract(r'\:(.*?)\-')
这里用原始字符串r''包裹正则,避免转义符出问题,(.*?)会精准捕获冒号和连字符之间的内容。
方法2:用apply()配合re.findall()
如果非要用re.findall(),就得给每个字符串单独处理,通过apply()遍历Series的每个元素:
import re def get_target_num(s): matches = re.findall(r'\:(.*?)\-', s) return matches[0] if matches else None all_cancers['提取的数值'] = all_cancers.iloc[:,3].astype(str).apply(get_target_num)
验证结果
处理完后查看新列:
print(all_cancers['提取的数值'])
会得到:
0 100414771 1 10506157 2 109655506 3 113903257 4 117598869 Name: 提取的数值, dtype: object
内容的提问来源于stack exchange,提问作者Achal Neupane
相关产品推荐
相关产品推荐

