Python处理大型CSV文件:全球区域人口数据分类问题求助
Hey there! Let’s work through your large CSV population data processing project step by step. First, I see your code cuts off at root.geome...—let’s fix that gap and build out the core functionality for reading, parsing, and writing those age-grouped population records.
First, let’s finish the basic Tkinter window and add a button to select your target CSV file. This will let you pick large files without hardcoding paths:
from tkinter import * from tkinter import filedialog import csv root = Tk() root.title("Population Data Processor") root.geometry("400x150") # Set a reasonable window size def select_csv_file(): file_path = filedialog.askopenfilename( filetypes=[("CSV Files", "*.csv")], title="Select Population Data CSV" ) if file_path: process_population_data(file_path) # Add a button to trigger file selection select_btn = Button(root, text="Select CSV File", command=select_csv_file, padx=10, pady=5) select_btn.pack(pady=50) root.mainloop()
Now let’s build the process_population_data function to handle reading, parsing the three age groups, and writing processed data. Since you’re dealing with large CSVs, we’ll read rows one at a time to avoid loading the entire file into memory:
def process_population_data(file_path): # Define output file path (you can adjust this as needed) output_path = "processed_population_data.csv" with open(file_path, mode='r', newline='', encoding='utf-8') as input_file, \ open(output_path, mode='w', newline='', encoding='utf-8') as output_file: # Use DictReader to access columns by name (adjust column names to match your CSV) reader = csv.DictReader(input_file) # Define output columns (add/remove based on your needs) fieldnames = ["Region", "1-15 Population", "15-65 Population", "65+ Population", "Total Population"] writer = csv.DictWriter(output_file, fieldnames=fieldnames) writer.writeheader() for row in reader: # Parse each age group (replace column names with your actual CSV headers) try: age_1_15 = int(row["1-15岁人口"]) age_15_65 = int(row["15-65岁人口"]) age_65_plus = int(row["65岁以上人口"]) except ValueError: print(f"Warning: Invalid numeric value in row {reader.line_num}") continue # Skip problematic rows or handle them as needed total = age_1_15 + age_15_65 + age_65_plus # Write processed row to output CSV writer.writerow({ "Region": row["区域"], "1-15 Population": age_1_15, "15-65 Population": age_15_65, "65+ Population": age_65_plus, "Total Population": total }) print(f"Processing complete! Output saved to {output_path}")
- Memory Efficiency: Using
csv.DictReaderand iterating row-by-row ensures we don’t load the entire large CSV into RAM, which prevents crashes or slowdowns. - Error Resilience: The added try-except block catches non-numeric population values, so your script won’t crash unexpectedly if there’s messy data in your CSV.
- Encoding Flexibility: If you run into encoding errors, swap
encoding='utf-8'withencoding='latin-1'—it’s a common fallback for CSVs with non-standard character sets.
If you’re hitting specific issues (like column name mismatches, unexpected data formats, or slow processing), share more details about your CSV structure or the exact error messages you’re seeing, and we can tweak this further!
内容的提问来源于stack exchange,提问作者Tann Mann

