You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为DataFrame新增District列:基于邮编所属区域列表的问题咨询

Fixing the ValueError & Better Pandas Approaches

First, let’s address that ValueError you’re seeing: it occurs because your result list ends up shorter than your codes DataFrame. If any zip code in main_zips isn’t present in d2, d5, or d7, it never gets added to result—so the lengths don’t match. Adding an else clause to append a default value (like None or "Unknown") would fix the error, but there are far cleaner, faster ways to handle this in pandas without loops.

Here are two optimized approaches:

1. Use a Mapping Dictionary (Fast & Concise)

Create a single dictionary that maps each zip code to its district, then use pandas’ map() method to assign the district values in one line. This is vectorized (way faster than loops for large datasets) and easy to read.

# Build the district-zip mapping
district_mapping = {}
for district_name in ['District 2', 'District 5', 'District 7']:
    # Drop NaN values in case some columns have empty entries
    zip_codes = d_codes[district_name].dropna().tolist()
    for zip_code in zip_codes:
        district_mapping[zip_code] = district_name

# Assign the District column to codes
codes['District'] = codes['Zip Code'].map(district_mapping)

If a zip code doesn’t match any district, it’ll get a NaN value—you can fill these with a default using codes['District'] = codes['District'].fillna("Unknown") if needed.

2. Reshape & Merge (Clean for Tabular Data)

Use pandas’ melt() function to turn your wide d_codes DataFrame into a long format, then merge it with codes. This is great if you prefer working with tabular joins instead of dictionaries.

# Reshape d_codes from wide to long format
melted_districts = d_codes.melt(
    var_name='District', 
    value_name='Zip Code'
).dropna()  # Remove rows with missing zip codes

# Merge with codes to get the district for each zip
codes = codes.merge(melted_districts, on='Zip Code', how='left')

The how='left' ensures all rows from codes are kept, even if a zip code has no matching district (those will get NaN in the District column).

Why These Are Better Than Loops

  • Speed: Pandas vectorized operations (like map() and merge()) are implemented in C, so they’re orders of magnitude faster than Python for loops, especially with large datasets.
  • Readability: The code is more concise and clearly expresses your intent (mapping zips to districts, or joining related data).
  • Maintainability: If you add more districts later, you just need to update the list of district names instead of modifying a loop.

内容的提问来源于stack exchange,提问作者Zubair Lakhia

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 20:22:56