如何使用Python统计并打印CSV文件的总记录数?
Hey there! Let's work through how to get the total record count from your CSV file—sounds like you tried sum and count approaches without luck, so let's cover reliable methods that handle CSV structure properly.
Method 1: Use Python's Built-in csv Module (No Third-Party Libraries)
This is great if you want to stick to Python's standard library and avoid installing extra packages. It properly parses CSV rows, even if they contain line breaks within cells.
import csv # Replace with your actual CSV file path csv_path = "your_data.csv" total_records = 0 # Open the file with proper encoding and newline handling with open(csv_path, mode='r', newline='', encoding='utf-8') as csv_file: csv_reader = csv.reader(csv_file) # Uncomment this line if your CSV has a header row you want to exclude from the count # next(csv_reader) # Iterate through each row and increment the count for _ in csv_reader: total_records += 1 print(f"Total records: {total_records}")
Key Notes:
- The
withstatement ensures the file is closed automatically after processing. - Use
next(csv_reader)to skip the header row if your CSV has one (so you only count data rows). - This method correctly handles edge cases like cells with embedded line breaks, which raw line-counting methods would mess up.
Method 2: Use pandas (Simpler for Data Workflows)
If you're already working with data in Python, pandas makes this task super concise. First, install pandas if you haven't:
pip install pandas
Then use this code:
import pandas as pd csv_path = "your_data.csv" # Read the CSV into a DataFrame # Add `header=None` if your CSV doesn't have a header row df = pd.read_csv(csv_path) # Get the number of rows (records) total_records = df.shape[0] # Alternatively, `len(df)` works too! # total_records = len(df) print(f"Total records: {total_records}")
Why Your Previous sum/count Attempts Might Have Failed
Raw line-counting tricks like sum(1 for line in open(csv_path)) or open(csv_path).read().count('\n') often fail because:
- They count empty lines as records.
- They don't account for line breaks inside CSV cells (a valid CSV can have cells with newlines, which these methods would treat as separate rows).
- If the last line of your CSV doesn't end with a newline, you'll undercount by 1.
Using the csv module or pandas avoids these issues because they properly parse the CSV structure according to the format's rules.
内容的提问来源于stack exchange,提问作者rohan

