如何使用对应值列表填充DataFrame中的空单元格?
Hey there! Let's walk through how to fill those empty cells in your DataFrame. Here are a couple of straightforward approaches you can use, depending on your preference:
Approach 1: Use Dictionary Mappings
First, we'll create mapping dictionaries that link each city (with missing values) to its corresponding state and country. Then we'll target the empty cells in the original DataFrame and fill them using these mappings.
Step-by-Step Code:
import pandas as pd # Create your original DataFrame data = { 'City': ['Chicago', 'Boston', 'San Diego', 'Los Angeles', 'San Francisco', 'Sacramento', 'Vancouver', 'Toronto'], 'State': ['IL', '', 'CA', 'CA', '', '', 'BC', ''], 'Country': ['United States', '', 'United States', 'United States', '', '', 'Canada', ''] } df = pd.DataFrame(data) # Your pre-defined fill lists city_list = ['Boston', 'San Francisco', 'Sacramento', 'Toronto'] state_list = ['MA', 'CA', 'CA', 'ON'] country_list = ['United States', 'United States', 'United States', 'Canada'] # Create mapping dictionaries state_mapping = dict(zip(city_list, state_list)) country_mapping = dict(zip(city_list, country_list)) # Fill empty State cells df.loc[df['State'] == '', 'State'] = df.loc[df['State'] == '', 'City'].map(state_mapping) # Fill empty Country cells df.loc[df['Country'] == '', 'Country'] = df.loc[df['Country'] == '', 'City'].map(country_mapping)
Approach 2: Use a Fill DataFrame with combine_first
Another clean way is to create a separate DataFrame from your fill lists, then merge it with the original to fill missing values. This method works well if you want to handle all columns at once.
Step-by-Step Code:
import pandas as pd # Original DataFrame (same as above) data = { 'City': ['Chicago', 'Boston', 'San Diego', 'Los Angeles', 'San Francisco', 'Sacramento', 'Vancouver', 'Toronto'], 'State': ['IL', '', 'CA', 'CA', '', '', 'BC', ''], 'Country': ['United States', '', 'United States', 'United States', '', '', 'Canada', ''] } df = pd.DataFrame(data) # Create fill DataFrame from your lists fill_data = pd.DataFrame({ 'City': ['Boston', 'San Francisco', 'Sacramento', 'Toronto'], 'State': ['MA', 'CA', 'CA', 'ON'], 'Country': ['United States', 'United States', 'United States', 'Canada'] }) # Set City as index for both DataFrames, combine, then reset index df = df.set_index('City').combine_first(fill_data.set_index('City')).reset_index()
Resulting DataFrame
After running either approach, your DataFrame will look like this:
| City | State | Country |
|---|---|---|
| Chicago | IL | United States |
| Boston | MA | United States |
| San Diego | CA | United States |
| Los Angeles | CA | United States |
| San Francisco | CA | United States |
| Sacramento | CA | United States |
| Vancouver | BC | Canada |
| Toronto | ON | Canada |
Note: If your empty cells are NaN instead of empty strings, replace df['State'] == '' with pd.isnull(df['State']) in Approach 1.
内容的提问来源于stack exchange,提问作者Karma

