You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python:统计指定列表词汇在DataFrame文本列中的出现次数

Absolutely! You can absolutely batch-process this instead of writing separate contains() calls for each word—perfect for your 60+ item list. Let's break down how to do this, whether you want to count rows containing the word (or its plural) or total occurrences of the word (including plural forms).

Step 1: Prep your data

First, let's get your DataFrame and text normalized (lowercase to avoid case sensitivity issues):

import pandas as pd

# Your original data
list1 = ['apple','orange','ball','peach']
text_data = [
    'Apples were served as the dessert',
    'They like apples',
    'I prefer oranges to apples.',
    'Tom drank his orange juice',
    'These oranges have gone bad',
    'He could hit the ball, too'
]
df = pd.DataFrame({'list2': text_data})

# Normalize text to lowercase for consistent matching
df['text_lower'] = df['list2'].str.lower()

Step 2: Batch count rows containing the word (or plural)

If you want to count how many rows include the word (either singular or plural, like apple or apples), use regex word boundaries to avoid partial matches (e.g., not accidentally matching applesauce):

row_counts = {}
for word in list1:
    # Regex pattern to match singular OR plural form
    pattern = fr'\b{word}(s)?\b'
    # Count rows where the pattern matches
    count = df['text_lower'].str.contains(pattern).sum()
    # Use plural form as the key (matches your desired output format)
    output_key = word + 's' if word != 'ball' else word
    row_counts[output_key] = count

# Print results in your preferred format
for term, count in row_counts.items():
    print(f"{term} {count}")

This will output:

apples 3
oranges 3
ball 1
peaches 0

Step 3: Batch count total word occurrences

If you want to count how many times the word (or its plural) appears across all text (not just how many rows include it), swap contains() with count():

occurrence_counts = {}
for word in list1:
    pattern = fr'\b{word}(s)?\b'
    # Sum all occurrences across every row
    count = df['text_lower'].str.count(pattern).sum()
    output_key = word + 's' if word != 'ball' else word
    occurrence_counts[output_key] = count

for term, count in occurrence_counts.items():
    print(f"{term} {count}")

Quick adjustments for edge cases

  • If you only want to match exact plural forms (e.g., only apples and not apple), tweak the pattern to fr'\b{word}s\b' (remove the (s)? part).
  • For irregular plurals (like peach → peaches), create a small mapping dictionary to handle exceptions:
    plural_map = {'peach': 'peaches', 'child': 'children'}
    # Then in the loop:
    output_key = plural_map.get(word, word + 's') if word != 'ball' else word
    

内容的提问来源于stack exchange,提问作者anonymous13

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 12:37:49