You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于CSV数据,在Pandas中筛选DataFrame30-300范围值至另一DataFrame

Solution to Filter Values in 30-300 Range into New DataFrame

Got it, let's work through this problem step by step using pandas. The goal is to scan each row of your CSV, pull out values from numeric columns (mean, min, max, std) that fall between 30 and 300, then organize those values into a new DataFrame with their original column names. We'll also handle edge cases like inf and missing values to avoid errors.

Here's the complete implementation:

Step 1: Import Libraries and Load Data

First, we'll import pandas and numpy (to handle infinite values), then load your CSV file:

import pandas as pd
import numpy as np

# Load the CSV into a DataFrame
df = pd.read_csv('your_data.csv')

# Replace infinite values and empty cells with NaN (to prevent filtering issues)
df = df.replace([np.inf, -np.inf], np.nan)

Step 2: Define a Row-Wise Filter Function

We'll create a function that processes each row, checks the numeric columns, and returns only the values that meet our 30-300 criteria:

def filter_row_values(row):
    # Specify which columns are numeric (we'll skip 'date' and 'metric')
    numeric_columns = ['mean', 'min', 'max', 'std']
    
    # Use a dictionary comprehension to filter values in the target range
    filtered_entries = {
        col: val 
        for col, val in row[numeric_columns].items() 
        if pd.notna(val) and 30 < val < 300
    }
    
    # Return as a Series to align with DataFrame columns automatically
    return pd.Series(filtered_entries)

Step 3: Apply the Filter and Build the Final DataFrame

Now we'll apply the function to every row, then combine the filtered results with the original date and metric columns to keep context:

# Apply the filter to each row
filtered_values_df = df.apply(filter_row_values, axis=1)

# Merge with original date/metric columns to maintain context for each entry
final_df = pd.concat([df[['date', 'metric']], filtered_values_df], axis=1)

Step 4: Check the Result

If you print final_df, you'll get a clean DataFrame where each row contains only the values from the original numeric columns that fall in the 30-300 range. Here's what it looks like:

datemetricmaxminstdmean
2018-03-15cpu34.0NaNNaNNaN
2018-03-16mem40.090.0NaNNaN
2018-03-17cpuNaNNaN143.22NaN
2018-03-18cpuNaNNaNNaN52.86
2018-03-20mem67.9645.33NaNNaN
2018-03-22cpuNaNNaN119.05NaN

Key Notes:

  • We convert inf and missing values to NaN so they don't break our range check.
  • The final DataFrame keeps all original numeric column names; cells where no value met the criteria are filled with NaN.
  • Retaining date and metric ensures you know exactly which measurement each filtered value belongs to.

内容的提问来源于stack exchange,提问作者Souvik Ray

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 08:02:52