Python:统计指定列表词汇在DataFrame文本列中的出现次数
Absolutely! You can absolutely batch-process this instead of writing separate contains() calls for each word—perfect for your 60+ item list. Let's break down how to do this, whether you want to count rows containing the word (or its plural) or total occurrences of the word (including plural forms).
Step 1: Prep your data
First, let's get your DataFrame and text normalized (lowercase to avoid case sensitivity issues):
import pandas as pd # Your original data list1 = ['apple','orange','ball','peach'] text_data = [ 'Apples were served as the dessert', 'They like apples', 'I prefer oranges to apples.', 'Tom drank his orange juice', 'These oranges have gone bad', 'He could hit the ball, too' ] df = pd.DataFrame({'list2': text_data}) # Normalize text to lowercase for consistent matching df['text_lower'] = df['list2'].str.lower()
Step 2: Batch count rows containing the word (or plural)
If you want to count how many rows include the word (either singular or plural, like apple or apples), use regex word boundaries to avoid partial matches (e.g., not accidentally matching applesauce):
row_counts = {} for word in list1: # Regex pattern to match singular OR plural form pattern = fr'\b{word}(s)?\b' # Count rows where the pattern matches count = df['text_lower'].str.contains(pattern).sum() # Use plural form as the key (matches your desired output format) output_key = word + 's' if word != 'ball' else word row_counts[output_key] = count # Print results in your preferred format for term, count in row_counts.items(): print(f"{term} {count}")
This will output:
apples 3 oranges 3 ball 1 peaches 0
Step 3: Batch count total word occurrences
If you want to count how many times the word (or its plural) appears across all text (not just how many rows include it), swap contains() with count():
occurrence_counts = {} for word in list1: pattern = fr'\b{word}(s)?\b' # Sum all occurrences across every row count = df['text_lower'].str.count(pattern).sum() output_key = word + 's' if word != 'ball' else word occurrence_counts[output_key] = count for term, count in occurrence_counts.items(): print(f"{term} {count}")
Quick adjustments for edge cases
- If you only want to match exact plural forms (e.g., only
applesand notapple), tweak the pattern tofr'\b{word}s\b'(remove the(s)?part). - For irregular plurals (like
peach→peaches), create a small mapping dictionary to handle exceptions:plural_map = {'peach': 'peaches', 'child': 'children'} # Then in the loop: output_key = plural_map.get(word, word + 's') if word != 'ball' else word
内容的提问来源于stack exchange,提问作者anonymous13

