基于词与组元组映射计算词频及给定DataFrame数据处理需求
Got it, let's figure out how to compute the total frequency of each word from your utterance DataFrame. Here's a straightforward approach using pandas, which is perfect for this kind of text frequency analysis.
The Problem Breakdown
You have a list of user phrases and how often each phrase was used. We need to count how many times each individual word appears across all these phrases, where each occurrence of a word in a phrase contributes the phrase's frequency to the word's total count. For example, if "Show me directions to Starbucks" was used 754 times, each word in that phrase gets +754 to their total.
Step-by-Step Solution
First, let's set up our sample data (you can replace this with your actual DataFrame):
import pandas as pd # Your original data data = { 'Utterance': [ 'Directions to Starbucks', 'Show me directions to Starbucks', 'Give me directions to Starbucks', 'Navigate me to Starbucks', 'Display navigation to Starbucks', 'Direct me to Starbucks', 'Navigate to Starbucks' ], 'Frequency': [1045, 754, 612, 498, 376, 201, 180] } df = pd.DataFrame(data)
1. Split Utterances into Words and Expand Rows
We'll split each utterance into individual words, then "explode" those lists into separate rows so each word is paired with the original phrase's frequency:
# Split each utterance into words and expand into rows df_words = df.assign(Word=df['Utterance'].str.split()).explode('Word')
2. Normalize Words (Optional but Recommended)
To avoid counting "Directions" and "directions" as separate words, convert all words to lowercase:
df_words['Word'] = df_words['Word'].str.lower()
3. Calculate Total Word Frequency
Now group by each word and sum up the frequencies to get the total count:
# Aggregate frequencies per word word_frequency = df_words.groupby('Word')['Frequency'].sum().reset_index() # Sort by highest frequency first for readability word_frequency = word_frequency.sort_values('Frequency', ascending=False) print(word_frequency)
What This Does
- Splitting & Exploding: The
str.split()turns each phrase into a list of words, andexplode()creates a new row for every word in that list. So if a phrase has 4 words, it becomes 4 rows each with the same frequency value. - Lowercasing: Ensures case insensitivity, which is almost always what you want for word frequency analysis.
- Grouping & Summing: This step adds up all the frequency values for each unique word, giving you the total number of times that word was mentioned across all phrases (weighted by how often each phrase was used).
Sample Output
Running the code will give you this sorted result:
| Word | Frequency |
|---|---|
| to | 3666 |
| starbucks | 3666 |
| directions | 2411 |
| me | 2065 |
| navigate | 678 |
| show | 754 |
| give | 612 |
| display | 376 |
| navigation | 376 |
| direct | 201 |
That's it! This approach is efficient and easy to adjust if you need to handle edge cases like punctuation (just add a step to clean the utterances first, e.g., removing commas or periods).
内容的提问来源于stack exchange,提问作者user_seaweed

