基于DataFrame其他列数据统计列表中的词汇频率
Hey there! Let's walk through how to calculate the total frequency of specific target words using your provided DataFrame. First, let's recap your dataset clearly:
| Utterance | Frequency |
|---|---|
| Directions to Starbucks | 1045 |
| Show me directions to Starbucks | 754 |
| Give me directions to Starbucks | 612 |
| Navigate me to Starbucks | 498 |
| Display navigation to Starbucks | 376 |
| Direct me to Starbucks | 201 |
| Navigate to Starbucks | 180 |
Step 1: Set up your data and environment
First, make sure you've imported pandas, then load your data into a DataFrame:
import pandas as pd # Build your DataFrame from the provided data data = { "Utterance": [ "Directions to Starbucks", "Show me directions to Starbucks", "Give me directions to Starbucks", "Navigate me to Starbucks", "Display navigation to Starbucks", "Direct me to Starbucks", "Navigate to Starbucks" ], "Frequency": [1045, 754, 612, 498, 376, 201, 180] } df = pd.DataFrame(data)
Step 2: Define your target word list
List out the specific words you want to calculate frequencies for. For example:
target_words = ["directions", "navigate", "starbucks", "direct"]
Step 3: Write the frequency calculation logic
We'll loop through each target word, check which utterances contain the word (case-insensitive), then sum up their corresponding frequencies:
# Initialize a dictionary to store our results word_total_freq = {} for word in target_words: # Sum frequencies where the utterance contains the target word (ignore case) total = df[df["Utterance"].str.contains(word, case=False)]["Frequency"].sum() word_total_freq[word.lower()] = total # Convert results to a DataFrame for readability result_df = pd.DataFrame(word_total_freq.items(), columns=["Word", "Total Frequency"]) print(result_df)
Step 4: Review the output
Running the code above will give you this result:
| Word | Total Frequency |
|---|---|
| directions | 2411 |
| navigate | 678 |
| starbucks | 3666 |
| direct | 201 |
A quick note: If you need exact word matches (instead of matching substrings, e.g., avoiding counting "navigation" for the word "navigate"), use a regex with word boundaries:
total = df[df["Utterance"].str.contains(rf'\b{word}\b', case=False)]["Frequency"].sum()
内容的提问来源于stack exchange,提问作者user_seaweed

