基于条件处理DataFrame:列值判断与数据插入操作
Hey there! Let's work through this pandas problem together based on your requirements. I'll break down each step with code examples that you can adapt easily:
Step 1: Set up the DataFrame
First, let's recreate your input DataFrame to work with:
import pandas as pd import numpy as np data = { 'a': [2, 5, 7], 'b': [33, 33, 33], 'c': [17, 17, 17], 'd': [0, 0, 0], 'e': [418, 415, 413], 'f': [-5, -5, -5], 'g': [-81, -116, -116], 'j': [np.nan, np.nan, np.nan], 'k': [14336, 14336, 14336], 'l': [81, 81, 81], 'm': [1, 0, 1], 'n': [462, 487, 462], 'o': [-24, -5, -24], 'p': [np.nan, 116, np.nan], 'q': [81, 81, 81], 'r': [1, 1, 1], 's': [462, 462, 462], 't': [-24, -24, -111] } df = pd.DataFrame(data)
Step 2: Implement Requirement 1
Check if column c equals 0, and if so, assign values from d, e, f to new columns x1, x2, x3:
# Initialize new columns with NaN (default value when condition isn't met) df[['x1', 'x2', 'x3']] = np.nan # Use .loc to safely assign values based on the condition df.loc[df['c'] == 0, ['x1', 'x2', 'x3']] = df.loc[df['c'] == 0, ['d', 'e', 'f']].values
Note: In your sample data, column c has values of 17, so no rows will trigger this assignment right now. If you update rows where c=0, this logic will kick in.
Step 3: Implement Requirement 2
Check if column l equals 0, and if so, assign values from m, n, o to new columns y1, y2, y3:
# Initialize new columns with NaN df[['y1', 'y2', 'y3']] = np.nan # Assign values where l == 0 df.loc[df['l'] == 0, ['y1', 'y2', 'y3']] = df.loc[df['l'] == 0, ['m', 'n', 'o']].values
Again, your sample data has l=81 for all rows, so this won't populate yet—but it's ready for when you have rows with l=0.
Step 4: Handle Requirement 3 (Partial Implementation)
You mentioned when column l equals 1, we need to check values in o and t, but the requirement wasn't fully specified. Here's a flexible framework you can adapt to your exact needs:
# Example: Add a column to flag if both o and t are negative when l=1 df['l1_check'] = np.nan df.loc[df['l'] == 1, 'l1_check'] = df.loc[df['l'] == 1].apply( lambda row: "Both o and t are negative" if (row['o'] < 0 and row['t'] < 0) else "Condition not met", axis=1 )
Just replace the lambda function's logic with whatever specific check you need (e.g., comparing values, calculating a metric, etc.). If you share the full details of what you need to do when l=1, I can refine this further!
内容的提问来源于stack exchange,提问作者Snowfire777

