You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何对比两个Pandas DataFrame的列名与数据类型并忽略形状差异?

Solution for Comparing Pandas DataFrame Column Names and Data Types

Let's break down how to get detailed differences between two DataFrames' schemas (column names + data types) and fix the assert_frame_equal shape mismatch issue.

1. Custom Function to Print Detailed Schema Differences

Instead of just checking if the dtype dictionaries are equal, we can build a function that explicitly calls out missing columns and data type mismatches:

import pandas as pd
import numpy as np

def compare_df_schema(df1, df2):
    # Check for column name differences
    df1_unique_cols = set(df1.columns) - set(df2.columns)
    df2_unique_cols = set(df2.columns) - set(df1.columns)
    
    if df1_unique_cols:
        print(f"Columns only present in df1: {sorted(df1_unique_cols)}")
    if df2_unique_cols:
        print(f"Columns only present in df2: {sorted(df2_unique_cols)}")
    
    # Check data type mismatches in shared columns
    shared_cols = set(df1.columns) & set(df2.columns)
    dtype_errors = []
    for col in shared_cols:
        dtype1 = df1[col].dtype
        dtype2 = df2[col].dtype
        if dtype1 != dtype2:
            dtype_errors.append(
                f"Column '{col}': df1 uses {dtype1}, df2 uses {dtype2}"
            )
    
    if dtype_errors:
        print("\nData type mismatches in shared columns:")
        for error in dtype_errors:
            print(f"  • {error}")
    else:
        print("\nAll shared columns have matching data types.")
    
    # Final schema similarity check
    schema_matches = not df1_unique_cols and not df2_unique_cols and not dtype_errors
    print(f"\nOverall schema similarity: {schema_matches}")

Example Usage

Using your sample data:

# Sample DataFrames
df1 = pd.DataFrame({'A': ['a', 'b'], 'B': ['c', 'd'], 'C': ['e', 'f']})
df2 = pd.DataFrame({'A': [1, 2], 'B': ['c', 'd'], 'C': ['e', 'f']})

compare_df_schema(df1, df2)

This will output:

All shared columns have matching data types.

Data type mismatches in shared columns:
  • Column 'A': df1 uses object, df2 uses int64

Overall schema similarity: False

2. Fixing assert_frame_equal Shape Mismatch

The assert_frame_equal function throws an error if the DataFrames have different shapes (different number of rows/columns). To work around this, first align the DataFrames to only include shared columns, then run the assertion:

def assert_df_schema_match(df1, df2):
    shared_cols = sorted(set(df1.columns) & set(df2.columns))
    if not shared_cols:
        raise AssertionError("No common columns between DataFrames")
    
    # Align both DataFrames to shared columns
    df1_aligned = df1[shared_cols]
    df2_aligned = df2[shared_cols]
    
    # Run assertion with check_like to ignore column order
    try:
        pd.testing.assert_frame_equal(df1_aligned, df2_aligned, check_like=True, check_dtype=True)
        print("Aligned DataFrames are identical (values and dtypes match)")
    except AssertionError as e:
        print("Aligned DataFrames do not match:")
        print(e)

Key Notes:

  • check_like=True: Ignores column order (so columns don't need to be in the same sequence)
  • check_dtype=True: Ensures data types are compared (default behavior, included here for clarity)
  • If you only care about schema (not values), you can use check_data=False to skip value comparisons

Final Thoughts

The custom compare_df_schema function gives you clear, actionable feedback about exactly what's different between your DataFrames' schemas. For formal assertions, aligning the DataFrames first avoids shape mismatch errors in assert_frame_equal.

内容的提问来源于stack exchange,提问作者user3447653

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 15:22:28