You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

遍历DataFrame并提取值:如何按缺失值占比生成等级列?

Efficient Solutions for Your Pandas DataFrame Tasks

Let's tackle your two requirements step by step, using pandas' built-in tools that are both efficient and readable—avoiding unnecessary loops whenever possible.

1. Traversing a DataFrame and Extracting Values

First off: you rarely need to use explicit for loops with pandas since it's designed for vectorized operations, which are way faster. That said, here are the best approaches depending on what you need:

Extracting Entire Columns (Fastest)

If you just need all values from a specific column, access it directly—this is the most efficient method:

# Extract all 'id' values as a Series
id_values = df['id']

# Convert to a list if needed
id_list = df['id'].tolist()

Iterating Over Rows (When Necessary)

If you need to process row-by-row, use df.itertuples() (faster than df.iterrows()) or df.iterrows():

# Using itertuples (preferred for speed)
for row in df.itertuples(index=False):
    print(f"ID: {row.id}, Missing Value: {row.missing_value}")

# Using iterrows (preserves index, slightly slower)
for idx, row in df.iterrows():
    print(f"Index {idx}: ID={row['id']}, Missing Value={row['missing value']}")

2. Adding a Missing Value Grade Column

For your ranking logic, pd.cut() is the perfect tool—it's built exactly for binning values into discrete intervals, and it's way cleaner than nested np.where statements.

Step-by-Step Implementation

Your rules are:

  • ≤25 → Grade 1
  • ≤50 → Grade 2 (covers 25 < value ≤50)
  • ≤75 → Grade 3 (covers 50 < value ≤75)
  • ≤80 → Grade 4 (covers 75 < value ≤80)

Here's how to implement this:

import pandas as pd

# Your sample DataFrame
df = pd.DataFrame({
    'id': ['1245', '1323', '1784', '1557','1456'],
    'value': [11558522, 12323552, 13770958, 18412280, 13770958],
    'missing value': [34, 56, 80, 5, 76]
})

# Define bins and corresponding labels
bins = [-float('inf'), 25, 50, 75, 80]
labels = [1, 2, 3, 4]

# Create the new 'grade' column
df['grade'] = pd.cut(df['missing value'], bins=bins, labels=labels, include_lowest=True)

# Check the result
print(df)

Output Explanation

Running this code will produce:

id      value  missing value grade
0  1245  11558522             34     2
1  1323  12323552             56     3
2  1784  13770958             80     4
3  1557  18412280              5     1
4  1456  13770958             76     4
  • include_lowest=True ensures values exactly equal to the first bin edge (25) are included in Grade 1.
  • The bins are set to cover every possible value below 80, with the first bin starting at negative infinity to catch any values ≤25.

Why This Is Better Than Nested np.where

If you tried using np.where, you'd end up with messy code like this:

import numpy as np

df['grade'] = np.where(df['missing value'] <=25, 1,
                       np.where(df['missing value'] <=50, 2,
                                np.where(df['missing value'] <=75,3,
                                         np.where(df['missing value'] <=80,4, np.nan))))

This works, but it's harder to read and maintain—especially if you ever need to adjust the bins or add new grades. pd.cut() keeps your logic clean and scalable.


内容的提问来源于stack exchange,提问作者Tamarie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 19:22:45