使用Python提取分号分隔文本中指定关键词后的目标单词
Extract Target Words After "thin" and "fat"
Approach
Let's break down how to pull the exact words you need from that text. The key here is using regular expressions to zero in on the segments following "thin" or "fat", then splitting those segments to grab individual words. We'll also clean up any extra whitespace to make sure we only get the valid words you want.
Solution Code (Python)
import re input_text = "Jet pack(human, metal) thin wire, sheet; fat tube,rod thin girl; fat boy; We like to read They like to write End" # Capture all content after "thin " or "fat " up to the next semicolon target_segments = re.findall(r'(?:thin|fat)\s+([^;]+)', input_text) # Process each segment to extract clean words final_words = [] for segment in target_segments: # Split by commas, strip whitespace, and skip empty strings cleaned_words = [word.strip() for word in segment.split(',') if word.strip()] final_words.extend(cleaned_words) print(final_words)
Output
['wire', 'sheet', 'tube', 'rod', 'girl', 'boy']
Explanation
- Regex Breakdown: The pattern
(?:thin|fat)\s+([^;]+)does two main things:(?:thin|fat): Matches either "thin" or "fat" without capturing the word itself\s+: Skips the space immediately after "thin" or "fat"([^;]+): Captures everything until the next semicolon (since semicolons separate different "thin"/"fat" groups)
- Cleaning Segments: Each captured segment is split by commas. We strip any extra whitespace from each word and filter out empty entries that might come from odd spacing or trailing commas.
- Final Result: The resulting list is exactly the set of words you can use as parameters for your function.
内容的提问来源于stack exchange,提问作者Hamad Hassan
相关产品推荐
相关产品推荐

