为DataFrame新增District列:基于邮编所属区域列表的问题咨询
First, let’s address that ValueError you’re seeing: it occurs because your result list ends up shorter than your codes DataFrame. If any zip code in main_zips isn’t present in d2, d5, or d7, it never gets added to result—so the lengths don’t match. Adding an else clause to append a default value (like None or "Unknown") would fix the error, but there are far cleaner, faster ways to handle this in pandas without loops.
Here are two optimized approaches:
1. Use a Mapping Dictionary (Fast & Concise)
Create a single dictionary that maps each zip code to its district, then use pandas’ map() method to assign the district values in one line. This is vectorized (way faster than loops for large datasets) and easy to read.
# Build the district-zip mapping district_mapping = {} for district_name in ['District 2', 'District 5', 'District 7']: # Drop NaN values in case some columns have empty entries zip_codes = d_codes[district_name].dropna().tolist() for zip_code in zip_codes: district_mapping[zip_code] = district_name # Assign the District column to codes codes['District'] = codes['Zip Code'].map(district_mapping)
If a zip code doesn’t match any district, it’ll get a NaN value—you can fill these with a default using codes['District'] = codes['District'].fillna("Unknown") if needed.
2. Reshape & Merge (Clean for Tabular Data)
Use pandas’ melt() function to turn your wide d_codes DataFrame into a long format, then merge it with codes. This is great if you prefer working with tabular joins instead of dictionaries.
# Reshape d_codes from wide to long format melted_districts = d_codes.melt( var_name='District', value_name='Zip Code' ).dropna() # Remove rows with missing zip codes # Merge with codes to get the district for each zip codes = codes.merge(melted_districts, on='Zip Code', how='left')
The how='left' ensures all rows from codes are kept, even if a zip code has no matching district (those will get NaN in the District column).
Why These Are Better Than Loops
- Speed: Pandas vectorized operations (like
map()andmerge()) are implemented in C, so they’re orders of magnitude faster than Python for loops, especially with large datasets. - Readability: The code is more concise and clearly expresses your intent (mapping zips to districts, or joining related data).
- Maintainability: If you add more districts later, you just need to update the list of district names instead of modifying a loop.
内容的提问来源于stack exchange,提问作者Zubair Lakhia

