如何高效使用参考CSV更新目标CSV的国家对应大洲字段?
I see the issue with your current approach—you're building a dictionary that maps continents to lists of countries, but what we actually need is the reverse: a country-to-continent mapping. That's why your map() call is returning NaNs for most entries; it's trying to look up country names in keys that are continent names.
Let's fix this step by step, and also make the solution scalable for multiple reference CSVs.
Step 1: Correctly Build the Country-to-Continent Dictionary
First, we'll process the reference CSV to create a dictionary where each key is a country, and the value is its corresponding continent. We'll ignore any NaN values since they don't represent valid countries.
import pandas as pd # Load reference data continent_data = pd.read_csv('Continents.csv') # Initialize empty dictionary for country-to-continent mapping country_to_continent = {} # Iterate over each row in the reference data for _, row in continent_data.iterrows(): continent = row['Continent'] # Iterate over all country columns (skip the first 'Continent' column) for country in row[1:]: # Only add non-NaN entries to the dictionary if pd.notna(country): country_to_continent[country] = continent print(country_to_continent) # Output: {'U.S.': 'North America', 'Mexico': 'North America', 'Guatemala': 'North America', 'Honduras': 'North America', 'India': 'Asia', 'China': 'Asia', 'Nepal': 'Asia', 'Australia': 'Australia', 'South Africa': 'Africa', 'Botswana': 'Africa', 'Zimbabwe': 'Africa'}
Step 2: Apply the Mapping to Your Animals DataFrame
Now use this dictionary to replace the Country column in your Animals data:
# Load animals data animals_df = pd.read_csv('Animals.csv') # Replace Country with Continent using our mapping animals_df['Country'] = animals_df['Country'].map(country_to_continent) print(animals_df)
This will produce your desired output:
Name Species Country 0 Hobbes Tiger North America 1 SherKhan Tiger Asia 2 Rescuer Mouse Australia 3 Mickey Mouse North America
Step 3: Scalable Version for Multiple Reference CSVs
If you need to handle multiple reference CSVs, we can wrap the dictionary-building logic into a reusable function. This function will load all reference files, combine their mappings, and return a single country-to-continent dictionary:
def build_continent_mapping(reference_csv_paths): country_to_continent = {} for path in reference_csv_paths: df = pd.read_csv(path) for _, row in df.iterrows(): continent = row['Continent'] for country in row[1:]: if pd.notna(country): country_to_continent[country] = continent return country_to_continent # Example usage with multiple CSVs reference_files = ['Continents.csv', 'MoreContinents.csv'] combined_mapping = build_continent_mapping(reference_files) # Apply to animals data animals_df['Country'] = animals_df['Country'].map(combined_mapping)
This function will merge mappings from all provided CSVs. Note that if a country appears in multiple reference files, the last occurrence will overwrite previous entries—adjust this logic if you need to handle conflicts differently.
Why Your Original Code Failed
Your code transposed the reference data and created a dictionary like {'North America': ['U.S.', 'Mexico', ...], ...}. When you tried to map() this to the Country column, Pandas looked for country names (like 'U.S.') as keys in the dictionary, which don't exist—hence the NaNs. The reverse mapping fixes this entirely.
内容的提问来源于stack exchange,提问作者TheMagicRaj

