DataFrame多列内容比对及人性化可读格式转换技术咨询
Got it, let's break down your two pandas problems with practical, easy-to-follow solutions—these are super common when working with messy or unstructured tabular data!
1. Compare Multiple Columns in a DataFrame & Output Boolean Results to a New DataFrame
If you need to check equality (or other conditions) between pairs of columns and store those True/False results in a separate DataFrame, pandas makes this straightforward.
Example Scenario
Suppose your original DataFrame looks like this:
import pandas as pd original_df = pd.DataFrame({ 'Product': ['Laptop', 'Phone', 'Tablet', 'Headphones'], 'Price_2023': [999, 699, 299, 199], 'Price_2024': [999, 749, 299, 249] })
Solution Code
To compare Price_2023 vs Price_2024 and store the boolean results in a new DataFrame:
# Create a new DataFrame for boolean comparisons boolean_results = pd.DataFrame() # Add comparison columns—you can extend this to any number of column pairs boolean_results['Price_Unchanged'] = original_df['Price_2023'] == original_df['Price_2024'] boolean_results['Price_Increased'] = original_df['Price_2024'] > original_df['Price_2023'] # If you want to keep the original index or add reference columns, merge them together boolean_results = pd.concat([original_df[['Product']], boolean_results], axis=1)
Customization Tip
If you need more complex comparisons (like partial string matches), use pandas' string methods:
# Example: Check if Product name contains "Phone" (case-insensitive) boolean_results['Is_Phone'] = original_df['Product'].str.contains('Phone', case=False)
2. Restructure a DataFrame for Human Readability (Combine Duplicate Rows)
Your second task is to fix a DataFrame where the same question appears across multiple rows with different categories. We can use groupby() to collapse these duplicates into a single row per question, with all related categories grouped together.
Example Scenario
Your original DF1 looks like this:
df1 = pd.DataFrame({ 'column1': ['question 1', 'question 2', 'question 1'], 'column2': ['category A', 'category A', 'category B'], 'column3': ['subcategory A', 'subcategory B', 'subcategory C'] })
Solution Code
We'll group by the question column and aggregate the category columns into either comma-separated strings (for quick reading) or lists (for further processing):
# Option 1: Combine categories into comma-separated strings (great for human scanning) readable_df = df1.groupby('column1').agg({ 'column2': lambda x: ', '.join(x.unique()), # Use unique() to avoid duplicate categories 'column3': lambda x: ', '.join(x.unique()) }).reset_index() # Option 2: Combine categories into lists (better for downstream data operations) readable_df_lists = df1.groupby('column1').agg({ 'column2': list, 'column3': list }).reset_index()
Result Explanation
After running Option 1, your readable DataFrame will look like this:
| column1 | column2 | column3 |
|---|---|---|
| question 1 | category A, category B | subcategory A, subcategory C |
| question 2 | category A | subcategory B |
This format is way easier to scan at a glance!
内容的提问来源于stack exchange,提问作者ben

