Pandas str.extract循环正则提取数值时覆盖已有结果的问题排查
问题根因
循环中直接执行df['watt'] = 提取结果的写法会覆盖全列:str.extract返回的结果仅包含当前watt列为空的行的索引,全列赋值时,pandas会将不在该返回结果索引内的行(即之前已经匹配成功的行)自动填充为NaN,最终导致之前的提取结果丢失。
另外原代码存在基础疏漏:使用了np.nan但没有导入numpy库。
修复方法
方法1:用fillna实现增量填充(写法最简洁)
fillna只会对列中原有NaN的位置做填充,不会覆盖已经存在的非空值,完全匹配“保留之前提取结果”的需求:
import pandas as pd import numpy as np df = pd.DataFrame([{'title':'This bulb operates at 222 watts and is fabulous.'}, {'title':'This bulb operates at 999 w and is fantastic.'}]) # 正则前加r标记为原始字符串,避免转义异常 regexes = [r'([0-9\.,]{1,})[\s\-]{0,1}watt[s]{0,1} ', r'([0-9\.,]{1,})[\s\-]{0,1}w '] # 初始化空列 df['watt'] = np.nan for regex in regexes: # 仅填充空值,不覆盖已有结果 # expand=False让extract直接返回Series,和列结构对齐,避免赋值警告 df['watt'] = df['watt'].fillna(df['title'].str.extract(regex, expand=False)) print(df)
方法2:用.loc精准定位空行赋值
通过布尔索引只选中需要更新的行做赋值,完全不会触碰已经匹配成功的行:
import pandas as pd import numpy as np df = pd.DataFrame([{'title':'This bulb operates at 222 watts and is fabulous.'}, {'title':'This bulb operates at 999 w and is fantastic.'}]) regexes = [r'([0-9\.,]{1,})[\s\-]{0,1}watt[s]{0,1} ', r'([0-9\.,]{1,})[\s\-]{0,1}w '] df['watt'] = np.nan for regex in regexes: # 定位所有watt为空的行 null_mask = df['watt'].isnull() # 仅对空行做提取和赋值 df.loc[null_mask, 'watt'] = df.loc[null_mask, 'title'].str.extract(regex, expand=False) print(df)
运行结果
两种写法最终输出一致,两个数值都能正确提取,不会出现覆盖问题:
title watt 0 This bulb operates at 222 watts and is fabulous. 222 1 This bulb operates at 999 w and is fantastic. 999
内容的提问来源于stack exchange,提问作者Jabb
相关产品推荐
相关产品推荐

