You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

DataFrame多列内容比对及人性化可读格式转换技术咨询

Solution for Your Two Pandas DataFrame Tasks

Got it, let's break down your two pandas problems with practical, easy-to-follow solutions—these are super common when working with messy or unstructured tabular data!

1. Compare Multiple Columns in a DataFrame & Output Boolean Results to a New DataFrame

If you need to check equality (or other conditions) between pairs of columns and store those True/False results in a separate DataFrame, pandas makes this straightforward.

Example Scenario

Suppose your original DataFrame looks like this:

import pandas as pd

original_df = pd.DataFrame({
    'Product': ['Laptop', 'Phone', 'Tablet', 'Headphones'],
    'Price_2023': [999, 699, 299, 199],
    'Price_2024': [999, 749, 299, 249]
})

Solution Code

To compare Price_2023 vs Price_2024 and store the boolean results in a new DataFrame:

# Create a new DataFrame for boolean comparisons
boolean_results = pd.DataFrame()

# Add comparison columns—you can extend this to any number of column pairs
boolean_results['Price_Unchanged'] = original_df['Price_2023'] == original_df['Price_2024']
boolean_results['Price_Increased'] = original_df['Price_2024'] > original_df['Price_2023']

# If you want to keep the original index or add reference columns, merge them together
boolean_results = pd.concat([original_df[['Product']], boolean_results], axis=1)

Customization Tip

If you need more complex comparisons (like partial string matches), use pandas' string methods:

# Example: Check if Product name contains "Phone" (case-insensitive)
boolean_results['Is_Phone'] = original_df['Product'].str.contains('Phone', case=False)

2. Restructure a DataFrame for Human Readability (Combine Duplicate Rows)

Your second task is to fix a DataFrame where the same question appears across multiple rows with different categories. We can use groupby() to collapse these duplicates into a single row per question, with all related categories grouped together.

Example Scenario

Your original DF1 looks like this:

df1 = pd.DataFrame({
    'column1': ['question 1', 'question 2', 'question 1'],
    'column2': ['category A', 'category A', 'category B'],
    'column3': ['subcategory A', 'subcategory B', 'subcategory C']
})

Solution Code

We'll group by the question column and aggregate the category columns into either comma-separated strings (for quick reading) or lists (for further processing):

# Option 1: Combine categories into comma-separated strings (great for human scanning)
readable_df = df1.groupby('column1').agg({
    'column2': lambda x: ', '.join(x.unique()),  # Use unique() to avoid duplicate categories
    'column3': lambda x: ', '.join(x.unique())
}).reset_index()

# Option 2: Combine categories into lists (better for downstream data operations)
readable_df_lists = df1.groupby('column1').agg({
    'column2': list,
    'column3': list
}).reset_index()

Result Explanation

After running Option 1, your readable DataFrame will look like this:

column1column2column3
question 1category A, category Bsubcategory A, subcategory C
question 2category Asubcategory B

This format is way easier to scan at a glance!


内容的提问来源于stack exchange,提问作者ben

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:04:51