You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas DataFrame中基于Value列生成带后缀的最大数值新列Value2?

Solution to Extract Maximum Value with Suffix in Pandas

To solve your problem of creating the Value2 column, we'll use a combination of regex extraction and a helper function to parse and select the maximum value (preserving any suffix like %). Here's a step-by-step breakdown:

Step 1: Define the Helper Function

This function processes each entry in the Value column to meet your requirements:

  • Returns NaN if the entry is "--".
  • Uses regex to pull out all numeric segments (including optional % suffixes).
  • Compares the numeric values of these segments, then returns the one with the highest value (keeping its original format).
import pandas as pd
import re
import numpy as np

def extract_max_value(s):
    if s == "--":
        return np.nan
    # Extract all numeric parts (digits + optional %)
    numeric_segments = re.findall(r'(\d+%?)', s)
    if not numeric_segments:
        return np.nan
    # Convert segments to numeric values for comparison, while retaining original text
    parsed_segments = []
    for seg in numeric_segments:
        num = float(seg.replace('%', ''))
        parsed_segments.append((num, seg))
    # Pick the segment with the highest numeric value
    max_segment = max(parsed_segments, key=lambda x: x[0])[1]
    return max_segment

Step 2: Apply the Function to Your DataFrame

Let's test this with your sample data:

# Create your sample DataFrame
data = {
    'Id': [1, 2, 3, 4, 5],
    'Value': ['>45%', '>29%', '<30 to >69', '>40% to <56%', '--']
}
df = pd.DataFrame(data)

# Generate the Value2 column
df['Value2'] = df['Value'].apply(extract_max_value)

print(df)

Output

Running this code will produce exactly the desired result:

Id             Value Value2
0   1              >45%    45%
1   2              >29%    29%
2   3       <30 to >69     69
3   4  >40% to <56%   56%
4   5                --    NaN

How It Works

  • Regex Extraction: re.findall(r'(\d+%?)', s) captures every sequence of digits followed by an optional % (e.g., "45%" or "69").
  • Numeric Comparison: We temporarily strip % to convert segments to floats, find the maximum value, then return the original segment to keep the suffix intact.
  • Edge Case Handling: Directly checks for "--" to return NaN, and handles empty extractions gracefully.

内容的提问来源于stack exchange,提问作者James Lin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 20:47:42