数据集条目数量统计及指定词重复次数查询方法
Hey there! Let's walk through how to solve these two common data analysis tasks using pandas (a go-to tool for working with structured data in Python):
Getting the total number of rows (entries) in your dataset is straightforward. Once you've loaded your data into a pandas DataFrame, you have two simple options:
import pandas as pd # First, load your dataset (replace with your actual file path/loading method) df = pd.read_csv('your_data.csv') # Option 1: Use .shape[0] to get the row count total_entries = df.shape[0] # Option 2: Use len() on the DataFrame total_entries = len(df) print(f"Total entries in the dataset: {total_entries}")
Both methods work equally well—df.shape[0] returns the first value of the DataFrame's shape tuple (rows × columns), while len(df) directly counts the number of rows.
Let's say you have a column (like 'region') with values such as East, Northeast, West, and you want to count how many times "Northeast" appears. Here are two reliable ways to do this:
Exact match (case-sensitive by default)
# Option 1: Use value_counts() to get counts for all values, then fetch the target northeast_count = df['region'].value_counts().get('Northeast', 0) # Option 2: Use boolean masking with sum() to count matches northeast_count = (df['region'] == 'Northeast').sum() print(f"'Northeast' appears {northeast_count} times in the 'region' column")
The .get('Northeast', 0) ensures you get 0 instead of an error if "Northeast" doesn't exist in the column.
Bonus: Case-insensitive or partial match
If you need to count entries that contain "Northeast" (e.g., "East-Northeast") or match regardless of capitalization, use str.contains():
# Count all entries with "Northeast" (case-insensitive) northeast_related_count = df['region'].str.contains('Northeast', case=False).sum() print(f"Entries containing 'Northeast' (case-insensitive): {northeast_related_count}")
内容的提问来源于stack exchange,提问作者Justin Beshears

