You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python DataFrame按行统计精确单词次数及上下文提取报错问题

Got it, let's figure out what's going wrong here and fix it step by step!

问题分析

First off, the error you're seeing makes total sense—your original code df["string"].count(df["word"]) is using Pandas' Series.count() method incorrectly:

  • Series.count() doesn't count substring occurrences at all—it's designed to return the number of non-missing values in the series, which has nothing to do with string matching.
  • You're also passing the entire df["word"] series as an argument, and Pandas can't automatically map each row's string to its corresponding word this way.
解决方案

We need to process each row individually to handle two requirements: count exact matches of the word and extract 5 characters before and after each match. We'll use regular expressions to ensure exact matching (so apple won't be confused with apples) and Pandas' apply to handle row-by-row logic.

Full Code Example

import pandas as pd
import re

# Sample data matching your description
data = {
    "word": ["apple"],
    "string": [
        "One apple price is 40",
        "one apple price is 50 and three apples price are 60 but apple..."
    ]
}
df = pd.DataFrame(data)

def process_row(row):
    target_word = row["word"]
    text = row["string"]
    
    # Build regex for exact matching: word boundaries + escape special characters
    # Remove re.IGNORECASE if you need case-sensitive matching
    regex_pattern = re.compile(r"\b" + re.escape(target_word) + r"\b", re.IGNORECASE)
    
    # 1. Count exact matches
    match_count = len(regex_pattern.findall(text))
    
    # 2. Extract 5 chars before/after each match (handle edge cases)
    match_snippets = []
    for match in regex_pattern.finditer(text):
        # Avoid negative indices or out-of-bounds errors
        start_pos = max(0, match.start() - 5)
        end_pos = min(len(text), match.end() + 5)
        snippet = text[start_pos:end_pos]
        match_snippets.append(snippet)
    
    # Combine all snippets into a single summary string
    summary = "; ".join(match_snippets)
    
    return pd.Series([match_count, summary], index=["count", "wordSummary"])

# Apply the function to every row
df[["count", "wordSummary"]] = df.apply(process_row, axis=1)

print(df)

Key Details Explained

  • Exact Matching: The \b in the regex stands for word boundary, which ensures we only match standalone instances of apple (not parts of longer words like apples). re.escape(target_word) handles special characters in the word (like don't or $5) so they don't break the regex.
  • Case Sensitivity: If you need to distinguish between Apple and apple, just remove the re.IGNORECASE parameter.
  • Edge Case Handling: Using max(0, ...) and min(len(text), ...) prevents errors when the matched word is near the start or end of the string (e.g., no 5 characters available before it).
  • Efficiency: Precompiling the regex with re.compile speeds things up if you're working with a large dataset.

Sample Output

For your example data, the result will look like this:

word                                              string  count                                      wordSummary
0  apple                              One apple price is 40      1                               One apple price 
1  apple  one apple price is 50 and three apples price ar...      2  one apple price ; ut apple...

内容的提问来源于stack exchange,提问作者naqwe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 06:37:37