Python DataFrame按行统计精确单词次数及上下文提取报错问题
Got it, let's figure out what's going wrong here and fix it step by step!
问题分析
First off, the error you're seeing makes total sense—your original code df["string"].count(df["word"]) is using Pandas' Series.count() method incorrectly:
Series.count()doesn't count substring occurrences at all—it's designed to return the number of non-missing values in the series, which has nothing to do with string matching.- You're also passing the entire
df["word"]series as an argument, and Pandas can't automatically map each row'sstringto its correspondingwordthis way.
解决方案
We need to process each row individually to handle two requirements: count exact matches of the word and extract 5 characters before and after each match. We'll use regular expressions to ensure exact matching (so apple won't be confused with apples) and Pandas' apply to handle row-by-row logic.
Full Code Example
import pandas as pd import re # Sample data matching your description data = { "word": ["apple"], "string": [ "One apple price is 40", "one apple price is 50 and three apples price are 60 but apple..." ] } df = pd.DataFrame(data) def process_row(row): target_word = row["word"] text = row["string"] # Build regex for exact matching: word boundaries + escape special characters # Remove re.IGNORECASE if you need case-sensitive matching regex_pattern = re.compile(r"\b" + re.escape(target_word) + r"\b", re.IGNORECASE) # 1. Count exact matches match_count = len(regex_pattern.findall(text)) # 2. Extract 5 chars before/after each match (handle edge cases) match_snippets = [] for match in regex_pattern.finditer(text): # Avoid negative indices or out-of-bounds errors start_pos = max(0, match.start() - 5) end_pos = min(len(text), match.end() + 5) snippet = text[start_pos:end_pos] match_snippets.append(snippet) # Combine all snippets into a single summary string summary = "; ".join(match_snippets) return pd.Series([match_count, summary], index=["count", "wordSummary"]) # Apply the function to every row df[["count", "wordSummary"]] = df.apply(process_row, axis=1) print(df)
Key Details Explained
- Exact Matching: The
\bin the regex stands for word boundary, which ensures we only match standalone instances ofapple(not parts of longer words likeapples).re.escape(target_word)handles special characters in the word (likedon'tor$5) so they don't break the regex. - Case Sensitivity: If you need to distinguish between
Appleandapple, just remove there.IGNORECASEparameter. - Edge Case Handling: Using
max(0, ...)andmin(len(text), ...)prevents errors when the matched word is near the start or end of the string (e.g., no 5 characters available before it). - Efficiency: Precompiling the regex with
re.compilespeeds things up if you're working with a large dataset.
Sample Output
For your example data, the result will look like this:
word string count wordSummary 0 apple One apple price is 40 1 One apple price 1 apple one apple price is 50 and three apples price ar... 2 one apple price ; ut apple...
内容的提问来源于stack exchange,提问作者naqwe
相关产品推荐
相关产品推荐

