Python筛选DataFrame指定经纬度范围数据并导出CSV
Solution to Filter and Export Pandas DataFrame Rows by Latitude/Longitude
Hey there! Let's wrap up your task of filtering geographic data and exporting the results. You've already started with importing pandas—here's the complete, working code plus clear breakdowns:
Complete Working Code
import pandas as pd # Read the original CSV file into a DataFrame df = pd.read_csv("your_input_file.csv") # Replace with your actual file path # Filter rows where startlat is between 18-20 AND startlon is between 72.5-75 filtered_df = df[(df['startlat'].between(18, 20)) & (df['startlon'].between(72.5, 75))] # Export the filtered DataFrame to a new CSV file filtered_df.to_csv("filtered_geo_data.csv", index=False) # index=False removes the extra index column
Key Breakdowns
- Reading the CSV: Make sure to replace
"your_input_file.csv"with the actual path to your CSV file (relative or absolute). - Filtering Logic:
- We use
pd.Series.between()to cleanly check if values fall within the specified ranges—it's more readable than writing(df['startlat'] >=18) & (df['startlat'] <=20). - The
&operator ensures both conditions are met. Don't forget the parentheses around each condition—they're required for correct operator precedence!
- We use
- Exporting: The
index=Falseparameter is important because it prevents pandas from adding an extra column with the row index to your output CSV.
Optional Checks (To Avoid Headaches)
- Verify your column names match exactly: Run
print(df.columns)to double-check thatstartlatandstartlonare spelled correctly (pandas is case-sensitive!). - Handle missing values: If your dataset has NaN values in these columns, add
.dropna(subset=['startlat', 'startlon'])before filtering to exclude rows with missing coordinates:df_clean = df.dropna(subset=['startlat', 'startlon']) filtered_df = df_clean[(df_clean['startlat'].between(18,20)) & (df_clean['startlon'].between(72.5,75))]
内容的提问来源于stack exchange,提问作者Ajinkya
相关产品推荐
相关产品推荐

