You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python CSV文件ETL分析:行数、列数等统计功能异常排查

Troubleshooting Python CSV ETL Analysis Code (Pandas)

Hey there! Let's walk through your code and fix those issues step by step. You've got a solid start with pandas, but there are a few small missteps causing the unexpected behavior.

Key Issues in Your Current Code

Let's break down what's going wrong:

  1. Incorrect Argument Type for analysingData: Your function expects a pandas DataFrame (df) as input, but you're passing a file path string when you call it. Strings don't have attributes like shape or methods like isnull(), which is why you're seeing errors.
  2. Misaligned Empty Value Logic: Your loop prints null counts per row, but your requirements ask for column-level null/non-null stats and total nulls across all columns—this loop doesn't address those needs, and it's inefficient compared to pandas' built-in tools.
  3. Wasted Data Load: You load the DataFrame in main() but don't reuse it for analysis; instead, you pass a file path again, leading to redundant work and errors.

Fixed Code with Full ETL Analysis

import pandas as pd

def load_csv(file_path):
    # Load CSV and return the DataFrame (separates loading from analysis)
    return pd.read_csv(file_path)

def analysingData(df):
    # 1. Output total rows and columns
    print("Total number of rows: ", df.shape[0])
    print("Total number of columns: ", df.shape[1])
    print("\n")
    
    # 2. Output non-null rows per column
    print("Non-null rows per column:")
    print(df.count(axis=0))  # count() directly returns non-null values per column
    print("\n")
    
    # 3. Output null values per column
    print("Null values per column:")
    null_per_column = df.isnull().sum(axis=0)
    print(null_per_column)
    print("\n")
    
    # 4. Output total null values across all columns
    total_nulls = df.isnull().sum().sum()
    print(f"Total null values in all columns: {total_nulls}")
    print("\n")
    
    # 5. Output duplicate row count
    duplicate_count = df.duplicated(subset=None).sum()
    print(f"Number of duplicate rows: {duplicate_count}")
    print("\n")
    
    # Bonus: Optional view of rows with null values (only rows that have empty entries)
    print("Rows with null values (count per row):")
    row_null_counts = df.isnull().sum(axis=1)
    print(row_null_counts[row_null_counts > 0])

def main():
    file_path = r"C:\Users\aliceoc\.spyder-py3\ProgrammingforBigDataCA\fireAndAmbulance (3).csv"
    df = load_csv(file_path)
    analysingData(df)  # Pass the loaded DataFrame directly

if __name__ == "__main__":
    main()

What Changed & Why

  • Data Reuse: We load the CSV once in main() and pass the DataFrame to analysingData, eliminating redundant file reads and fixing the argument type error.
  • Pandas Built-ins: We replaced the row loop with efficient pandas methods:
    • df.count(axis=0): Exactly what you need for non-null column counts.
    • df.isnull().sum(axis=0): Calculates nulls per column in one line.
    • df.isnull().sum().sum(): Sums all column null counts to get the total across the dataset.
  • Clearer Output: Used f-strings to make stats more readable, and added a filtered view of rows with nulls (so you don't have to scroll through every row).

内容的提问来源于stack exchange,提问作者Cam13

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 07:28:18