Python 3.7下如何将CSV加载的DataFrame列格式化为整数
Hey there! Let's break down how to fix your two issues step by step: preserving integer formatting when loading CSVs, and identifying rows that exist in one DataFrame but not the other.
1. Fixing the Integer-to-Float Conversion Issue
The reason your integer values (like 2010392) are showing up as 2010392.0 is almost certainly because your CSV has missing values (NaN) in the IDDLECT column. Pandas defaults to using float dtype for columns with NaNs since standard integer types can't hold missing values.
Your earlier attempt with pd.to_numeric didn't work because you didn't assign the result back to the DataFrame column—Pandas operations return new objects by default, they don't modify the original data in place. Here's how to fix it properly:
Option 1: Specify Dtype When Loading CSV
Use Pandas' nullable integer type (Int64, capital I) when reading the CSV. This type supports both integers and NaNs, so your values stay in integer format without the .0 suffix:
import pandas as pd # Define dtype for the problematic column dtype_spec = {'IDDLECT': 'Int64'} data = pd.read_csv("timetable_all_2019-2_groups.csv", dtype=dtype_spec) data02 = data.drop_duplicates() # Verify the dtype and values print(f"IDDLECT column dtype: {data02['IDDLECT'].dtype}") # Will show 'Int64' print(data02['IDDLECT'].head(20)) # Values will be integers (e.g., 2010392, not 2010392.0)
Option 2: Convert Column After Loading
If you already loaded the data without specifying dtype, you can convert the column afterward—just remember to assign the result back:
# Convert the column and assign back to the DataFrame data02['IDDLECT'] = pd.to_numeric(data02['IDDLECT'], downcast='integer', errors='coerce').astype('Int64')
2. Finding Rows Unique to One DataFrame
Assuming you have a second DataFrame (let's call it data_other) and you want to find rows in data02 that don't exist in data_other, here are two reliable methods:
Method 1: Use merge with Indicator
This is the most straightforward approach, especially if you want to compare all columns:
# Merge the two DataFrames and keep track of which rows come from where merged = data02.merge(data_other, on=list(data02.columns), how='left', indicator=True) # Filter rows that only exist in data02 unique_rows = merged[merged['_merge'] == 'left_only'].drop(columns='_merge') print("Rows in data02 not present in data_other:") print(unique_rows)
Method 2: Use isin with Key Columns
If you only need to compare specific key columns (like IDDCYR, IDDSUBJ, IDDLECT) instead of all columns, this method is efficient:
# Define your unique identifier columns key_columns = ['IDDCYR', 'IDDSUBJ', 'IDDLECT'] # Create tuples of key values for each row, then check membership data02_keys = data02[key_columns].apply(tuple, axis=1) data_other_keys = data_other[key_columns].apply(tuple, axis=1) # Filter rows in data02 that aren't in data_other unique_rows = data02[~data02_keys.isin(data_other_keys)] print("Rows in data02 not present in data_other:") print(unique_rows)
Full Working Example
Putting it all together, here's a complete script that handles both tasks:
import pandas as pd # Load first CSV with correct dtype dtype_spec = {'IDDLECT': 'Int64'} data = pd.read_csv("timetable_all_2019-2_groups.csv", dtype=dtype_spec) data02 = data.drop_duplicates() print(f'Len data {len(data)}') print(data.head(20)) print(f'Len data02 {len(data02)}') print(data02.head(20)) # Load second CSV (adjust filename as needed) data_other = pd.read_csv("your_second_data.csv", dtype=dtype_spec) # Find rows unique to data02 merged = data02.merge(data_other, on=list(data02.columns), how='left', indicator=True) unique_to_data02 = merged[merged['_merge'] == 'left_only'].drop('_merge', axis=1) print("\nUnique rows in data02:") print(unique_to_data02)
内容的提问来源于stack exchange,提问作者Phlip

