You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效使用参考CSV更新目标CSV的国家对应大洲字段?

Solution to Map Countries to Continents Using Pandas

I see the issue with your current approach—you're building a dictionary that maps continents to lists of countries, but what we actually need is the reverse: a country-to-continent mapping. That's why your map() call is returning NaNs for most entries; it's trying to look up country names in keys that are continent names.

Let's fix this step by step, and also make the solution scalable for multiple reference CSVs.

Step 1: Correctly Build the Country-to-Continent Dictionary

First, we'll process the reference CSV to create a dictionary where each key is a country, and the value is its corresponding continent. We'll ignore any NaN values since they don't represent valid countries.

import pandas as pd

# Load reference data
continent_data = pd.read_csv('Continents.csv')

# Initialize empty dictionary for country-to-continent mapping
country_to_continent = {}

# Iterate over each row in the reference data
for _, row in continent_data.iterrows():
    continent = row['Continent']
    # Iterate over all country columns (skip the first 'Continent' column)
    for country in row[1:]:
        # Only add non-NaN entries to the dictionary
        if pd.notna(country):
            country_to_continent[country] = continent

print(country_to_continent)
# Output: {'U.S.': 'North America', 'Mexico': 'North America', 'Guatemala': 'North America', 'Honduras': 'North America', 'India': 'Asia', 'China': 'Asia', 'Nepal': 'Asia', 'Australia': 'Australia', 'South Africa': 'Africa', 'Botswana': 'Africa', 'Zimbabwe': 'Africa'}

Step 2: Apply the Mapping to Your Animals DataFrame

Now use this dictionary to replace the Country column in your Animals data:

# Load animals data
animals_df = pd.read_csv('Animals.csv')

# Replace Country with Continent using our mapping
animals_df['Country'] = animals_df['Country'].map(country_to_continent)

print(animals_df)

This will produce your desired output:

Name    Species         Country
0  Hobbes    Tiger  North America
1  SherKhan  Tiger            Asia
2  Rescuer  Mouse       Australia
3  Mickey    Mouse  North America

Step 3: Scalable Version for Multiple Reference CSVs

If you need to handle multiple reference CSVs, we can wrap the dictionary-building logic into a reusable function. This function will load all reference files, combine their mappings, and return a single country-to-continent dictionary:

def build_continent_mapping(reference_csv_paths):
    country_to_continent = {}
    for path in reference_csv_paths:
        df = pd.read_csv(path)
        for _, row in df.iterrows():
            continent = row['Continent']
            for country in row[1:]:
                if pd.notna(country):
                    country_to_continent[country] = continent
    return country_to_continent

# Example usage with multiple CSVs
reference_files = ['Continents.csv', 'MoreContinents.csv']
combined_mapping = build_continent_mapping(reference_files)

# Apply to animals data
animals_df['Country'] = animals_df['Country'].map(combined_mapping)

This function will merge mappings from all provided CSVs. Note that if a country appears in multiple reference files, the last occurrence will overwrite previous entries—adjust this logic if you need to handle conflicts differently.

Why Your Original Code Failed

Your code transposed the reference data and created a dictionary like {'North America': ['U.S.', 'Mexico', ...], ...}. When you tried to map() this to the Country column, Pandas looked for country names (like 'U.S.') as keys in the dictionary, which don't exist—hence the NaNs. The reverse mapping fixes this entirely.

内容的提问来源于stack exchange,提问作者TheMagicRaj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:25:35