如何在Pandas DataFrame中使用elif逻辑创建分类列
Hey there! Let's sort out that category labeling issue you're working through with your pandas DataFrame. The problem with your current if/elif/else approach is that when you run it directly on a pandas Series (like df['Sales']), you'll hit a ValueError—this is because the condition returns a boolean array (one True/False per row) instead of a single boolean value that standard if statements expect.
Pandas thrives on vectorized operations for these kinds of tasks, so here are three reliable solutions to get your desired Category column:
1. Use numpy.select() (Best for Multiple Conditions)
This method lets you define clear condition-value pairs and a default for all other cases, making it easy to read and scale if you add more categories later.
import numpy as np import pandas as pd # First, let's recreate your sample DataFrame (adjust if your actual data differs) df = pd.DataFrame({ 'Company': ['MC'] * 10, 'Sales': [360.0, 340.0, 338.5, 335.5, 235.0, 235.0, 234.0, 127.0, 121.0, 120.5] }) # Define your conditions and matching categories conditions = [ df['Sales'] >= 300, (df['Sales'] >= 200) & (df['Sales'] < 300) ] category_labels = ['Fast Mover', 'Medium Fast Mover'] # Assign the categories (default to 'Slow Mover' for all other cases) df['Category'] = np.select(conditions, category_labels, default='Slow Mover')
2. Use pandas.cut() (Perfect for Binning Numeric Data)
If your categories are based on continuous numeric ranges, pd.cut() is a clean, concise option. It automatically maps values to predefined bins:
df['Category'] = pd.cut( df['Sales'], bins=[-float('inf'), 200, 300, float('inf')], # Define your range boundaries labels=['Slow Mover', 'Medium Fast Mover', 'Fast Mover'] # Match bins to labels )
Note: pd.cut() uses left-closed, right-open intervals by default, so this setup correctly groups values <200 as Slow Mover, 200-299.999 as Medium Fast Mover, and 300+ as Fast Mover.
3. Use apply() (Simple for Small Datasets)
While not as efficient for large datasets (it processes rows one by one), a custom function with apply() works if you prefer the familiar if/elif structure:
def get_category(sales_value): if sales_value >= 300: return 'Fast Mover' elif 200 <= sales_value < 300: return 'Medium Fast Mover' else: return 'Slow Mover' df['Category'] = df['Sales'].apply(get_category)
Verify the Result
All three methods will produce your desired output:
| Company | Sales | Category |
|---|---|---|
| MC | 360.0 | Fast Mover |
| MC | 340.0 | Fast Mover |
| MC | 338.5 | Fast Mover |
| MC | 335.5 | Fast Mover |
| MC | 235.0 | Medium Fast Mover |
| MC | 235.0 | Medium Fast Mover |
| MC | 234.0 | Medium Fast Mover |
| MC | 127.0 | Slow Mover |
| MC | 121.0 | Slow Mover |
| MC | 120.5 | Slow Mover |
For most cases, I'd recommend numpy.select() or pandas.cut()—they're vectorized operations, so they'll run much faster on large datasets compared to apply().
内容的提问来源于stack exchange,提问作者Ahamed Moosa

