如何使用正则表达式去除字符串中多个空格后的所有字符
Hey there! Let's get that string trimming issue sorted out for you.
What's Wrong With Your Original Code?
Your regex (.*\s?)(\s{2,}.*) uses greedy matching (.*), which tries to consume as much of the string as possible. For example, when processing "pear sa", the first group .*\s? will gobble up "pear " (including that single trailing space), leaving only "sa" behind—which doesn't have two leading spaces to trigger the second group. That's why your code fails to modify those strings with multiple spaces after a single space.
Simple Solution: Use re.sub Directly
The easiest way to achieve your goal is to replace any sequence of two or more spaces (and everything after them) with an empty string. Here's how:
import re list_of_strings = ["apple", "orange ca", "pear sa", "banana sth"] final_list_of_strings = [re.sub(r'\s{2,}.*', '', s) for s in list_of_strings] print(final_list_of_strings) # Output: ["apple", "orange ca", "pear", "banana"]
Let's break down the regex:
\s{2,}: Matches 2 or more whitespace characters.*: Matches everything that comes after those spaces- Replacing this entire match with
''effectively removes all content starting from the first occurrence of two or more spaces.
Alternative: Using Match Groups (If You Prefer)
If you want to stick with capturing groups, switch to non-greedy matching (.*?) to avoid over-consuming the string. This ensures we stop at the first occurrence of two or more spaces:
import re regex_expression = r'^(.*?)(?:\s{2,}.*)?$' list_of_strings = ["apple", "orange ca", "pear sa", "banana sth"] final_list_of_strings = [] for name in list_of_strings: match = re.match(regex_expression, name) if match: final_list_of_strings.append(match.group(1)) print(final_list_of_strings) # Same desired output
Here's what this regex does:
^and$: Ensure we match the entire string.*?: Non-greedily matches characters until it hits either two+ spaces or the end of the string(?:\s{2,}.*)?: A non-capturing group (we don't need to extract this part) that matches two+ spaces and everything after; the?makes it optional so strings without multiple spaces are still matched.
Either approach will give you the exact result you're looking for!
内容的提问来源于stack exchange,提问作者cincin21

