You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas DataFrame列条件替换与列折叠处理问题求助

Problem: Convert Pandas DataFrame Columns to Lists of Index Values (Handling NaNs)

Original Data & Requirements

You have the following Pandas DataFrame with string-based 'NaN' values, indexed by Symbol:

import pandas as pd

fn1 = pd.DataFrame([['A', 'NaN', 'NaN', 9, 6], ['B', 'NaN', 2, 'NaN', 7], ['C', 3, 2, 'NaN', 10], ['D', 'NaN', 7, 'NaN', 'NaN'], ['E', 'NaN', 'NaN', 3, 3], ['F', 'NaN', 'NaN', 7,'NaN']], columns = ['Symbol', 'Condition1','Condition2', 'Condition3', 'Condition4'])
fn1.set_index('Symbol', inplace=True)

Which looks like this:

SymbolCondition1Condition2Condition3Condition4
ANaNNaN96
BNaN2NaN7
C32NaN10
DNaN7NaNNaN
ENaNNaN33
FNaNNaN7NaN

Your goal is to:

  1. Process each column: replace non-NaN values with their corresponding row's Symbol index
  2. Collapse each column into a list of Symbols that had valid values
  3. Create a new DataFrame that retains the original column names, with each column holding its respective Symbol list

The Issue You Faced

Your initial code generated a nested list of Symbols per column, but you couldn't build the target DataFrame because the list lengths are inconsistent.


Solution

Don't worry—Pandas fully supports lists of varying lengths as column elements. We can adjust your approach to build a dictionary first (mapping column names to their Symbol lists), then convert that dictionary directly into a DataFrame. Here's the step-by-step fix:

Step 1: Clean the Data (Convert String 'NaN' to Actual NaN)

First, we need to turn the string 'NaN' values into Pandas-recognizable missing values (pd.NA), otherwise numerical comparisons (like >0) will fail:

fn1 = fn1.replace('NaN', pd.NA)

Step 2: Build the Result Dictionary & Convert to DataFrame

Use a dictionary comprehension to iterate over each column, filter out missing values, and collect the corresponding Symbol indices. Then convert this dictionary to your target DataFrame:

# Create a dict where keys are column names, values are lists of Symbols with valid entries
result_dict = {col: fn1[col].dropna().index.tolist() for col in fn1.columns}

# Convert the dict to a DataFrame
result_df = pd.DataFrame([result_dict])

What This Does

  • fn1[col].dropna() removes all rows where the column has missing values
  • .index.tolist() extracts the Symbol indices of those valid rows into a list
  • Wrapping the dictionary in [] when creating the DataFrame ensures each column holds the full list (instead of expanding the list into multiple rows)

Final Result

Running this code gives you the desired DataFrame:

Condition1Condition2Condition3Condition4
['C']['B', 'C', 'D']['A', 'E', 'F']['A', 'B', 'C', 'E']

If you ever want to expand these lists into individual rows (one Symbol per row), you can use the explode() method:

expanded_df = result_df.explode(list(result_df.columns))

内容的提问来源于stack exchange,提问作者user987443

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 09:08:10