如何在Python中将DataFrame相似类别列的字符串映射生成新列
Got it, let's tackle this problem. You want to group similar voice commands in your coffee_directions_df DataFrame by standardizing the utterances and creating new mapped columns. Here are two straightforward approaches using pandas:
Approach 1: Extract Coffee Shop Names & Standardize Utterances
This method uses regular expressions to identify the coffee shop in each utterance, then builds a standardized command format.
Step 1: Import pandas and load your data
First, make sure you have pandas imported (and define your DataFrame if you haven't already):
import pandas as pd # Your existing DataFrame coffee_directions_df = pd.DataFrame({ 'Utterance': [ 'Directions to Starbucks', 'Directions to Tullys', 'Give me directions to Tullys', 'Directions to Seattles Best', 'Show me directions to Dunkin', 'Directions to Daily Dozen', 'Show me directions to Starbucks', 'Give me directions to Dunkin', 'Navigate me to Seattles Best', 'Display navigation to Starbucks', 'Direct me to Starbucks' ], 'Frequency': [1045, 1034, 986, 875, 812, 789, 754, 612, 498, 376, 201] })
Step 2: Define coffee shops and extract them
We'll use a regex pattern to match any of the coffee shop names in the utterances:
# List all unique coffee shops from your data coffee_shops = ['Starbucks', 'Tullys', 'Seattles Best', 'Dunkin', 'Daily Dozen'] # Build a regex pattern to match any of these shops shop_pattern = '|'.join(coffee_shops) # Extract the coffee shop name into a new column coffee_directions_df['Coffee_Shop'] = coffee_directions_df['Utterance'].str.extract(f'({shop_pattern})', expand=False) # Create a standardized utterance column (unify all commands to "Directions to [Shop]") coffee_directions_df['Standardized_Utterance'] = 'Directions to ' + coffee_directions_df['Coffee_Shop']
Approach 2: Batch Replace with Regex Mapping
If you prefer a direct replacement approach, you can define a mapping of pattern-to-standardized-text:
# Define a regex map: match any utterance containing the shop, replace with a standard command standardize_map = { r'.*Starbucks': 'Directions to Starbucks', r'.*Tullys': 'Directions to Tullys', r'.*Seattles Best': 'Directions to Seattles Best', r'.*Dunkin': 'Directions to Dunkin', r'.*Daily Dozen': 'Directions to Daily Dozen' } # Apply the mapping to create the standardized column coffee_directions_df['Standardized_Utterance'] = coffee_directions_df['Utterance'].replace(standardize_map, regex=True) # Extract the coffee shop name from the standardized text coffee_directions_df['Coffee_Shop'] = coffee_directions_df['Standardized_Utterance'].str.split(' to ').str[-1]
Result Example
After running either approach, your DataFrame will look like this (abbreviated):
| Utterance | Frequency | Coffee_Shop | Standardized_Utterance |
|---|---|---|---|
| Directions to Starbucks | 1045 | Starbucks | Directions to Starbucks |
| Show me directions to Starbucks | 754 | Starbucks | Directions to Starbucks |
| Direct me to Starbucks | 201 | Starbucks | Directions to Starbucks |
| Give me directions to Tullys | 986 | Tullys | Directions to Tullys |
This makes it easy to analyze aggregate metrics (like total frequency per coffee shop) or clean up your voice command data for further processing.
内容的提问来源于stack exchange,提问作者user_seaweed

