如何在pandas中提取列字符串包含的最高百分比数值
pandas提取字符串列最高百分比数值实现方案
核心逻辑是通过正则匹配提取所有百分比数值,转换为数字后取最大值,以下是可直接运行的实现:
完整代码
import pandas as pd import re # 构造示例数据,替换为你自己的DataFrame即可 df = pd.DataFrame({ 'target_col': [ 'XX: (+2, 30%); (-5, 20%); (+17, 50%)', 'XX: (+8, 45%); (-3, 12%); (+21, 38%)' ] }) def extract_max_percent(raw_str): # 匹配所有百分比对应的数字部分 percent_numbers = re.findall(r'(\d+)%', raw_str) # 转换为整数后取最大值,无匹配时返回空值避免报错 return max(map(int, percent_numbers)) if percent_numbers else pd.NA # 应用函数生成新列存储最高百分比 df['max_percent'] = df['target_col'].apply(extract_max_percent)
运行后示例输出的max_percent列值分别为50、45,符合需求。
适配场景调整
- 若你的数据中存在小数百分比(如
12.5%),将正则改为r'(\d+\.?\d*)%',同时把int替换为float即可适配。 - 若需要最终结果保留
%符号,修改返回值为f"{max(map(int, percent_numbers))}%"即可。
内容的提问来源于stack exchange,提问作者Ran Antes
相关产品推荐
相关产品推荐

