遍历DataFrame并提取值:如何按缺失值占比生成等级列?
Let's tackle your two requirements step by step, using pandas' built-in tools that are both efficient and readable—avoiding unnecessary loops whenever possible.
1. Traversing a DataFrame and Extracting Values
First off: you rarely need to use explicit for loops with pandas since it's designed for vectorized operations, which are way faster. That said, here are the best approaches depending on what you need:
Extracting Entire Columns (Fastest)
If you just need all values from a specific column, access it directly—this is the most efficient method:
# Extract all 'id' values as a Series id_values = df['id'] # Convert to a list if needed id_list = df['id'].tolist()
Iterating Over Rows (When Necessary)
If you need to process row-by-row, use df.itertuples() (faster than df.iterrows()) or df.iterrows():
# Using itertuples (preferred for speed) for row in df.itertuples(index=False): print(f"ID: {row.id}, Missing Value: {row.missing_value}") # Using iterrows (preserves index, slightly slower) for idx, row in df.iterrows(): print(f"Index {idx}: ID={row['id']}, Missing Value={row['missing value']}")
2. Adding a Missing Value Grade Column
For your ranking logic, pd.cut() is the perfect tool—it's built exactly for binning values into discrete intervals, and it's way cleaner than nested np.where statements.
Step-by-Step Implementation
Your rules are:
- ≤25 → Grade 1
- ≤50 → Grade 2 (covers 25 < value ≤50)
- ≤75 → Grade 3 (covers 50 < value ≤75)
- ≤80 → Grade 4 (covers 75 < value ≤80)
Here's how to implement this:
import pandas as pd # Your sample DataFrame df = pd.DataFrame({ 'id': ['1245', '1323', '1784', '1557','1456'], 'value': [11558522, 12323552, 13770958, 18412280, 13770958], 'missing value': [34, 56, 80, 5, 76] }) # Define bins and corresponding labels bins = [-float('inf'), 25, 50, 75, 80] labels = [1, 2, 3, 4] # Create the new 'grade' column df['grade'] = pd.cut(df['missing value'], bins=bins, labels=labels, include_lowest=True) # Check the result print(df)
Output Explanation
Running this code will produce:
id value missing value grade 0 1245 11558522 34 2 1 1323 12323552 56 3 2 1784 13770958 80 4 3 1557 18412280 5 1 4 1456 13770958 76 4
include_lowest=Trueensures values exactly equal to the first bin edge (25) are included in Grade 1.- The bins are set to cover every possible value below 80, with the first bin starting at negative infinity to catch any values ≤25.
Why This Is Better Than Nested np.where
If you tried using np.where, you'd end up with messy code like this:
import numpy as np df['grade'] = np.where(df['missing value'] <=25, 1, np.where(df['missing value'] <=50, 2, np.where(df['missing value'] <=75,3, np.where(df['missing value'] <=80,4, np.nan))))
This works, but it's harder to read and maintain—especially if you ever need to adjust the bins or add new grades. pd.cut() keeps your logic clean and scalable.
内容的提问来源于stack exchange,提问作者Tamarie

