Python如何移除字符串开头以Hi或Hello起始的首个句子?
Hey there! Let's get your regex sorted so you can strip out those initial Hi/Hello sentences exactly how you want.
The Problem with Your Current Approach
Your current re.match is correctly finding the initial sentence, but you need to either replace that matched part with nothing or slice the string from the end of the match to get the remaining content. The key is to make sure your regex targets only sentences starting with Hi/Hello, and captures up to the end of that first sentence (including any trailing spaces).
Solution 1: Use re.sub for Clean Replacement
This is the most straightforward way—we'll write a regex that matches the unwanted initial sentence, then replace it with an empty string. For strings that don't start with Hi/Hello, they'll stay untouched.
Here's the code:
import re docs = ['Hi, my name is Eric. Are you blue?', "Hi, I'm ! What is your name?", 'This is a great idea. I would love to go.', 'Hello, I am Jane Brown. What is your name?', "Hello, I am a doctor! Let's go to the mall.", 'I am ready to go. Mom says hello.'] # Regex pattern to match Hi/Hello-started first sentence pattern = r'^(Hi|Hello).*?[.!?](?:\s+|$)' processed_docs = [re.sub(pattern, '', doc) for doc in docs] print(processed_docs)
Breaking Down the Regex
Let's unpack what each part does:
^: Anchors the match to the start of the string—ensures we only target the first sentence.(Hi|Hello): Matches either "Hi" or "Hello" as the starting word..*?: Non-greedy match for any characters—stops at the first sentence-ending punctuation instead of going all the way to the end of the string.[.!?]: Matches any standard sentence-ending punctuation (period, exclamation, question mark).(?:\s+|$): Non-capturing group that matches one or more spaces after the punctuation, or the end of the string. This ensures we don't leave a leading space in the remaining text.
Solution 2: Use re.match with String Slicing
If you prefer working with the match object directly, you can check if a match exists, then slice the string from the end of the match to get the remaining content:
import re docs = ['Hi, my name is Eric. Are you blue?', "Hi, I'm ! What is your name?", 'This is a great idea. I would love to go.', 'Hello, I am Jane Brown. What is your name?', "Hello, I am a doctor! Let's go to the mall.", 'I am ready to go. Mom says hello.'] pattern = r'^(Hi|Hello).*?[.!?](?:\s+|$)' processed_docs = [] for doc in docs: match = re.match(pattern, doc) if match: # Take everything after the matched initial sentence processed_docs.append(doc[match.end():]) else: # Keep the original string if no match processed_docs.append(doc) print(processed_docs)
Both Methods Give Your Desired Output
Either approach will produce exactly what you want:
['Are you blue?', 'What is your name?', 'This is a great idea. I would love to go.', 'What is your name?', "Let's go to the mall.", 'I am ready to go. Mom says hello.']
内容的提问来源于stack exchange,提问作者mjoy

