关于构建提取‘further’前后各两个单词逻辑的技术问询
提取词汇"further"前后各两个单词的实现方案
Hey there! Let's break down how to pull the two words before and after "further" in any sentence—even when there aren't enough words on either side (like when "further" is at the start or end of a sentence). I'll walk through the logic with concrete examples and a working code snippet.
核心思路
The plan is straightforward:
- Split the sentence into individual words, cleaning up any attached punctuation.
- Locate the position(s) of "further" (we'll handle uppercase "Further" too).
- Grab up to two words before it (if they exist) and up to two words after it (if they exist).
示例句子与预期结果
Let's use 5 example sentences to show how this works:
- Sentence 1:
We need to advance further
→ Words before: to, advance; Words after: 无 - Sentence 2:
Further one morning we met again
→ Words before: 无; Words after: one, morning - Sentence 3:
She decided to look further into the matter
→ Words before: to, look; Words after: into, the - Sentence 4:
After discussing this further with our team
→ Words before: discussing, this; Words after: with, our - Sentence 5:
To go further than expected is challenging
→ Words before: To, go; Words after: than, expected
Python 实现代码
Here's a robust Python function that handles edge cases, punctuation, and even multiple instances of "further" in one sentence:
import re def extract_further_context(sentence): # Split sentence into words, preserving apostrophes (e.g., "don't") and stripping punctuation words = re.findall(r"\b[\w']+\b", sentence) # Find all positions of "further" (case-insensitive match) further_positions = [i for i, word in enumerate(words) if word.lower() == "further"] # Collect results for each instance of "further" context_results = [] for pos in further_positions: # Get up to 2 words before (avoid negative indices) before_words = words[max(0, pos - 2) : pos] # Get up to 2 words after (avoid going beyond the word list) after_words = words[pos + 1 : min(pos + 3, len(words))] context_results.append({ "original_sentence": sentence, "words_before": before_words, "words_after": after_words }) return context_results # Test the function with our example sentences test_sentences = [ "We need to advance further", "Further one morning we met again", "She decided to look further into the matter", "After discussing this further with our team", "To go further than expected is challenging" ] for idx, sent in enumerate(test_sentences, 1): result = extract_further_context(sent)[0] print(f"Sentence {idx}: {sent}") print(f"Words before 'further': {', '.join(result['words_before']) if result['words_before'] else 'None'}") print(f"Words after 'further': {', '.join(result['words_after']) if result['words_after'] else 'None'}") print("---")
代码细节说明
- Punctuation Handling: The regex
\b[\w']+\bensures we don't treat "further." or "further," as part of the target word, and keeps words like "don't" intact. - Case Insensitivity: We check
word.lower() == "further"so both "further" and "Further" are matched. - Edge Case Protection: Using
max(0, pos-2)prevents errors when "further" is at the start of the sentence, andmin(pos+3, len(words))stops us from trying to access words that don't exist at the end. - Multiple Matches: If a sentence has more than one "further", the function returns context for each instance.
内容的提问来源于stack exchange,提问作者thrinadhn
相关产品推荐
相关产品推荐

