如何使用Python re模块提取文本中以元音字母开头的单词
元音开头单词提取实现方案
你之前想的「替换掉所有辅音开头单词」的思路能跑,但容错率低,处理文本里的换行、破折号、标点、转义引号时很容易出问题。更稳妥的方式是正向匹配直接提取所有符合要求的单词,不用处理各种杂七杂八的非目标内容。
核心正则逻辑说明
正则规则用r'\b[aeiouAEIOU][a-zA-Z]*\b',各部分作用:
\b:单词边界锚点,自动把引号、标点、换行、破折号这些非单词内容挡在匹配范围外,不会把"、逗号这类符号算进单词里[aeiouAEIOU]:匹配单词开头的元音字符,覆盖大小写场景(比如句首大写的A、小写的on都能命中)[a-zA-Z]*:匹配元音开头后跟着的所有英文字母,直到碰到非字母字符就停止- 搭配
re.findall()直接返回所有匹配到的单词列表,不用写复杂的替换逻辑、不用额外处理lambda表达式
可运行完整代码
import re # 定义原始文本 text_one = '"A Ukrainian American woman who lives near Boston, Massachusetts, told Fox News Digital on Monday that she can no longer speak on the phone with her own mother, who lives in southern Ukraine, because of the Russian attacks on Ukraine and the fear these attacks have engendered.' text_two = '\n\nMany people in southern Ukraine — as well as throughout the country — are right now living in fear for their lives as Russian soldiers overtake the area, the Boston-area woman said."' # 合并两段文本 full_content = text_one + text_two # 提取所有元音开头的单词 vowel_start_words = re.findall(r'\b[aeiouAEIOU][a-zA-Z]*\b', full_content) # 拼接成空格分隔的字符串输出 output = ' '.join(vowel_start_words) print(output)
实际运行输出
A Ukrainian American on on own in Ukraine of attacks on Ukraine and engendered in Ukraine as as throughout are in overtake area
补充:如果坚持要用替换逻辑的写法
如果你一定要按最初的替换思路实现,可以用下面的正则,但要注意这种写法容易残留多余空格,需要额外做空格清理:
# 替换所有辅音开头的单词+后面跟着的非字母字符,最后清理首尾多余空格 result = re.sub(r'\s*\b[^aeiouAEIOU\W][a-zA-Z]*\b[^a-zA-Z]*', ' ', full_content).strip()
内容的提问来源于stack exchange,提问作者ADWrobo
相关产品推荐
相关产品推荐

