正则表达式需求:匹配指定字符串并过滤含无关行的文本
Solution for Your Regular Expression Adjustment
Let's work through this adjusted requirement together. First, let's recap what you need:
- Completely ignore lines with only irrelevant text (like "bla bla")
- Pull out all valid matching strings from lines that do contain targets:
- Any string starting with
thban(case-insensitive, with zero or more non-whitespace characters after it — this coversTHBAN,THBANES900,THBANMP900,thbanes900, and all similar variants) - Any string starting with
C3followed by word characters (likeC3950,C3850)
- Any string starting with
- Combine all extracted strings into one space-separated output
Adjusted Regular Expression
To capture all your target strings while skipping irrelevant lines entirely, use this regex with global (g) and case-insensitive (i) modifiers:
\b(thban\S*|C3\w+)
How It Breaks Down
Let's unpack the pattern to make it clear:
\b: A word boundary, so we don't accidentally match partial words (like athbansubstring inside a longer unrelated word)thban\S*: Matchesthban(regardless of uppercase/lowercase) followed by 0 or more non-whitespace characters — this handles standaloneTHBANas well as longer variantsC3\w+: MatchesC3followed by 1 or more word characters (letters, numbers, underscores), which fits all yourC3-prefixed examples|: Alternation, so the regex matches either thethbanpattern or theC3pattern
How to Use It
- Run a global, case-insensitive match of this regex against your entire file content
- Collect all the matched results into a list
- Join the list with spaces, and you'll get exactly the output you want:
THBANES900 C3950 THBAN THBANES901 C3850 THBANMP900 thbanes900
For Find/Replace Tools (Text Editors, etc.)
If you're using a tool that supports regex replace instead of direct matching:
- Find:
[\s\S]*?(\b(thban\S*|C3\w+))|[\s\S]+ - Replace:
$1 - Finally, trim any leading/trailing spaces and collapse multiple spaces into one (most editors have a built-in option for this, or you can do a quick follow-up replace of
\s+with)
This works by either capturing a valid target string and keeping it, or matching irrelevant content (including entire lines with no targets) and replacing it with nothing except the captured target (plus a space to keep results separated).
内容的提问来源于stack exchange,提问作者Marco
相关产品推荐
相关产品推荐

