如何用re.findall编写正则:仅匹配RPD_PC_MISC,排除RPD_PC_Basic_Sanity_ST?
Solution to Match Only
RPD_PC_MISC (Exclude Similar Patterns) If you need to only extract RPD_PC_MISC from your JSON string while excluding matches like RPD_PC_Basic_Sanity_ST, you have a couple of straightforward regex options depending on your exact needs:
Option 1: Exact Full-Word Match (Most Common)
Use word boundaries (\b) to ensure you're matching the complete RPD_PC_MISC string—this prevents accidental partial hits like XRPD_PC_MISC or RPD_PC_MISCY:
import re json_content = ''' { "test1": "RPD_PC_MISC", "test2": "RPD_PC_Basic_Sanity_ST", "test3": "Random text with RPD_PC_MISC embedded", "test4": "RPD_PC_MISC_extra" } ''' # Regex pattern with word boundaries to enforce full-word match pattern = r'\bRPD_PC_MISC\b' matches = re.findall(pattern, json_content) print(matches) # Output: ['RPD_PC_MISC', 'RPD_PC_MISC']
Option 2: Exclude Suffix Underscores (If Needed)
If you also want to exclude cases where RPD_PC_MISC is followed by an underscore (like RPD_PC_MISC_extra), add a negative lookahead ((?!_)) to block unwanted extensions right after the match:
# Regex pattern to exclude underscore suffixes pattern = r'RPD_PC_MISC(?!_)' matches = re.findall(pattern, json_content) print(matches) # Output: ['RPD_PC_MISC', 'RPD_PC_MISC'] (excludes test4's entry)
Why This Works
- The original broad pattern
RPD_PC_\w+matches any word characters (including underscores) afterRPD_PC_, which is why it catches unwanted strings likeRPD_PC_Basic_Sanity_ST. - Our targeted patterns explicitly lock in
MISCas the only valid suffix: word boundaries ensure it's a standalone term, while the negative lookahead blocks any immediate underscore extensions.
内容的提问来源于stack exchange,提问作者om tripathi
相关产品推荐
相关产品推荐

