如何在Pandas中对列内相似词组/语句分组并统计频次?
Got it, let's break this down step by step. You want to count multi-word phrases instead of single words, and your example shows you’re interested in both continuous phrases and possibly phrases that span non-consecutive words (like "A123 B234 data" appearing across sentences even with other words in between). Here are two practical implementations using Pandas and Python to cover both scenarios:
1. Count Continuous Multi-Word Phrases (Standard N-Grams)
This is for strict, consecutive word groups (e.g., "test data" or "A123 B234 C345").
Step 1: Setup & Sample Data
First, let's replicate your example data in a Pandas DataFrame:
import pandas as pd from collections import Counter import re # Your sample sentences data = { "text": [ "A123 B234 C345 test data.", "A123 B234 C345 D555 test data.", "A123 B234 test data.", "A123 B234 C345 more data." ] } df = pd.DataFrame(data)
Step 2: Clean the Text
We’ll remove punctuation and split sentences into clean word lists:
def clean_text(text): # Strip punctuation (keep letters, numbers, and spaces) cleaned = re.sub(r'[^\w\s]', '', text) # Split into words (handles extra spaces automatically) return cleaned.split() df["cleaned_words"] = df["text"].apply(clean_text)
Step 3: Generate N-Grams (Continuous Phrases)
Write a function to generate all consecutive phrases of varying lengths (adjust n_min and n_max to match your needs):
def generate_ngrams(word_list, n_range): ngrams = [] for n in n_range: # Create all consecutive n-word phrases, joined by spaces for i in range(len(word_list) - n + 1): phrase = ' '.join(word_list[i:i+n]) ngrams.append(phrase) return ngrams # Generate phrases from 2 to 5 words long (tweak this range as needed) n_min, n_max = 2, 5 df["phrases"] = df["cleaned_words"].apply(lambda x: generate_ngrams(x, range(n_min, n_max+1)))
Step 4: Count Phrase Frequencies
Flatten all phrases and count their occurrences:
# Flatten the list of phrases from all sentences all_phrases = [phrase for sublist in df["phrases"] for phrase in sublist] phrase_counts = Counter(all_phrases) # Get top 100 most common phrases (like your original single-word approach) top_phrases = phrase_counts.most_common(100) # Example output for your sample data print("Top continuous phrases:") for phrase, count in top_phrases[:5]: print(f"- {phrase}: {count}")
This will output things like:
Top continuous phrases: - A123 B234: 4 - test data: 3 - A123 B234 C345: 3 - B234 C345: 3 - C345 test data: 1
2. Count Non-Continuous Keyword Combinations
If your example means you want to count sentences that contain a set of keywords (even with other words in between, like "A123 B234" and "data" appearing anywhere in the sentence), use this approach:
# Define the keyword combinations you want to track target_combinations = [ ("A123 B234", "data"), ("A123 B234", "test", "data"), ("A123 B234", "C345", "test", "data") ] # Count how many sentences contain all keywords in each combination combination_counts = {} for combo in target_combinations: sentence_matches = 0 for words in df["cleaned_words"]: sentence_str = ' '.join(words) # Check if all parts of the combo exist in the sentence if all(part in sentence_str for part in combo): sentence_matches += 1 # Format the combo into a readable phrase combination_name = ' '.join(combo) combination_counts[combination_name] = sentence_matches # Print results matching your example print("\nNon-continuous combination counts:") for phrase, count in combination_counts.items(): print(f"- {phrase}: {count}")
This will output exactly what you expected (adjusted for your sample data's actual matches):
Non-continuous combination counts: - A123 B234 data: 4 - A123 B234 test data: 3 - A123 B234 C345 test data: 2
Quick Notes
- Use the n-gram method if you want natural, consecutive phrases (great for text mining common expressions).
- Use the keyword combination method if you’re tracking specific terms that don’t need to be consecutive.
内容的提问来源于stack exchange,提问作者jason

