如何移除邮件正文每行末尾的统一格式日期时间戳?
It looks like your current approach is trying to filter out words that match date patterns or am/pm, but that's not targeting the actual issue—your datetime stamps are appended to the end of every line, not standalone words. Let's fix this with a regex that directly targets and removes those trailing timestamps.
The Correct Regex Approach
Your timestamps follow a consistent format: YYYY-MM-DD HH:MM:SS at the end of each line. We can use re.sub() with a multiline flag to target these at the end of every line and replace them with an empty string.
Working Code Example
import re # Your input email body body = """Dear Volunteer2018-05-21 19:59:15 Your booking has been updated at metrowitnessing.com .2018-05-21 19:59:15 Crown Street - June 15th, 10:00am2018-05-21 19:59:15 Anthony James (m: 04xxxxxxxx)2018-05-21 19:59:15 Monica Brown (m: 04xxxxxxxx)2018-05-21 19:59:15 Bob Smith (m: 04xxxxxxxx)2018-05-21 19:59:15 Status: Confirmed2018-05-21 19:59:15""" # Replace trailing datetime stamps on every line cleaned_content = re.sub(r'\d{4}-\d{2}-\d{2} \d{2}:\d{2}:\d{2}$', '', body, flags=re.MULTILINE) # Print each cleaned line for line in cleaned_content.splitlines(): print(line.strip()) # .strip() removes any accidental leading/trailing whitespace
What This Regex Does
\d{4}-\d{2}-\d{2}: Matches the date part (e.g.,2018-05-21)\d{2}:\d{2}:\d{2}: Matches the time part with a leading space (e.g.,19:59:15)$: Ensures we only match this pattern at the end of a lineflags=re.MULTILINE: Makes$apply to each line's end instead of just the end of the entire text
Output Result
After running the code, you'll get the clean content you want:
Dear Volunteer Your booking has been updated at metrowitnessing.com . Crown Street - June 15th, 10:00am Anthony James (m: 04xxxxxxxx) Monica Brown (m: 04xxxxxxxx) Bob Smith (m: 04xxxxxxxx) Status: Confirmed
Why Your Original Code Didn't Work
Your filter() approach uses re.match() to check if a word starts with a date or am/pm—but your timestamps are glued to the end of the actual content (not separate words). This means the regex never matches the right part of the string, so nothing gets filtered out. Replacing the trailing pattern directly is the right solution here.
内容的提问来源于stack exchange,提问作者Tim Butler

