Python CSV文件ETL分析:行数、列数等统计功能异常排查
Troubleshooting Python CSV ETL Analysis Code (Pandas)
Hey there! Let's walk through your code and fix those issues step by step. You've got a solid start with pandas, but there are a few small missteps causing the unexpected behavior.
Key Issues in Your Current Code
Let's break down what's going wrong:
- Incorrect Argument Type for
analysingData: Your function expects a pandas DataFrame (df) as input, but you're passing a file path string when you call it. Strings don't have attributes likeshapeor methods likeisnull(), which is why you're seeing errors. - Misaligned Empty Value Logic: Your loop prints null counts per row, but your requirements ask for column-level null/non-null stats and total nulls across all columns—this loop doesn't address those needs, and it's inefficient compared to pandas' built-in tools.
- Wasted Data Load: You load the DataFrame in
main()but don't reuse it for analysis; instead, you pass a file path again, leading to redundant work and errors.
Fixed Code with Full ETL Analysis
import pandas as pd def load_csv(file_path): # Load CSV and return the DataFrame (separates loading from analysis) return pd.read_csv(file_path) def analysingData(df): # 1. Output total rows and columns print("Total number of rows: ", df.shape[0]) print("Total number of columns: ", df.shape[1]) print("\n") # 2. Output non-null rows per column print("Non-null rows per column:") print(df.count(axis=0)) # count() directly returns non-null values per column print("\n") # 3. Output null values per column print("Null values per column:") null_per_column = df.isnull().sum(axis=0) print(null_per_column) print("\n") # 4. Output total null values across all columns total_nulls = df.isnull().sum().sum() print(f"Total null values in all columns: {total_nulls}") print("\n") # 5. Output duplicate row count duplicate_count = df.duplicated(subset=None).sum() print(f"Number of duplicate rows: {duplicate_count}") print("\n") # Bonus: Optional view of rows with null values (only rows that have empty entries) print("Rows with null values (count per row):") row_null_counts = df.isnull().sum(axis=1) print(row_null_counts[row_null_counts > 0]) def main(): file_path = r"C:\Users\aliceoc\.spyder-py3\ProgrammingforBigDataCA\fireAndAmbulance (3).csv" df = load_csv(file_path) analysingData(df) # Pass the loaded DataFrame directly if __name__ == "__main__": main()
What Changed & Why
- Data Reuse: We load the CSV once in
main()and pass the DataFrame toanalysingData, eliminating redundant file reads and fixing the argument type error. - Pandas Built-ins: We replaced the row loop with efficient pandas methods:
df.count(axis=0): Exactly what you need for non-null column counts.df.isnull().sum(axis=0): Calculates nulls per column in one line.df.isnull().sum().sum(): Sums all column null counts to get the total across the dataset.
- Clearer Output: Used f-strings to make stats more readable, and added a filtered view of rows with nulls (so you don't have to scroll through every row).
内容的提问来源于stack exchange,提问作者Cam13
相关产品推荐
相关产品推荐

