如何对比两个Pandas DataFrame的列名与数据类型并忽略形状差异?
Let's break down how to get detailed differences between two DataFrames' schemas (column names + data types) and fix the assert_frame_equal shape mismatch issue.
1. Custom Function to Print Detailed Schema Differences
Instead of just checking if the dtype dictionaries are equal, we can build a function that explicitly calls out missing columns and data type mismatches:
import pandas as pd import numpy as np def compare_df_schema(df1, df2): # Check for column name differences df1_unique_cols = set(df1.columns) - set(df2.columns) df2_unique_cols = set(df2.columns) - set(df1.columns) if df1_unique_cols: print(f"Columns only present in df1: {sorted(df1_unique_cols)}") if df2_unique_cols: print(f"Columns only present in df2: {sorted(df2_unique_cols)}") # Check data type mismatches in shared columns shared_cols = set(df1.columns) & set(df2.columns) dtype_errors = [] for col in shared_cols: dtype1 = df1[col].dtype dtype2 = df2[col].dtype if dtype1 != dtype2: dtype_errors.append( f"Column '{col}': df1 uses {dtype1}, df2 uses {dtype2}" ) if dtype_errors: print("\nData type mismatches in shared columns:") for error in dtype_errors: print(f" • {error}") else: print("\nAll shared columns have matching data types.") # Final schema similarity check schema_matches = not df1_unique_cols and not df2_unique_cols and not dtype_errors print(f"\nOverall schema similarity: {schema_matches}")
Example Usage
Using your sample data:
# Sample DataFrames df1 = pd.DataFrame({'A': ['a', 'b'], 'B': ['c', 'd'], 'C': ['e', 'f']}) df2 = pd.DataFrame({'A': [1, 2], 'B': ['c', 'd'], 'C': ['e', 'f']}) compare_df_schema(df1, df2)
This will output:
All shared columns have matching data types. Data type mismatches in shared columns: • Column 'A': df1 uses object, df2 uses int64 Overall schema similarity: False
2. Fixing assert_frame_equal Shape Mismatch
The assert_frame_equal function throws an error if the DataFrames have different shapes (different number of rows/columns). To work around this, first align the DataFrames to only include shared columns, then run the assertion:
def assert_df_schema_match(df1, df2): shared_cols = sorted(set(df1.columns) & set(df2.columns)) if not shared_cols: raise AssertionError("No common columns between DataFrames") # Align both DataFrames to shared columns df1_aligned = df1[shared_cols] df2_aligned = df2[shared_cols] # Run assertion with check_like to ignore column order try: pd.testing.assert_frame_equal(df1_aligned, df2_aligned, check_like=True, check_dtype=True) print("Aligned DataFrames are identical (values and dtypes match)") except AssertionError as e: print("Aligned DataFrames do not match:") print(e)
Key Notes:
check_like=True: Ignores column order (so columns don't need to be in the same sequence)check_dtype=True: Ensures data types are compared (default behavior, included here for clarity)- If you only care about schema (not values), you can use
check_data=Falseto skip value comparisons
Final Thoughts
The custom compare_df_schema function gives you clear, actionable feedback about exactly what's different between your DataFrames' schemas. For formal assertions, aligning the DataFrames first avoids shape mismatch errors in assert_frame_equal.
内容的提问来源于stack exchange,提问作者user3447653

