You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何让Python代码仅打印字母字符?正则匹配尝试未成功

问题与解决方案

问题描述

现有一段Python代码用于处理标准输入文本,规则是对每个字符输出:字符本身、固定值1、是否为单词起始元音(1是0否)、是否为单词结尾元音(1是0否)。但当前代码会输出非字母字符(比如撇号),期望处理类似"It's a beautiful life"的文本时,仅保留并输出字母字符,过滤掉所有特殊符号。

原始代码

import sys
import re

pattern = re.compile("^[a-z]+$")  # matches purely alphabetic words
starting_vowels = re.compile("(^[aeiouAEIOU])")  # matches starting vowels
ending_vowels = re.compile("[aeiouAEIOU]$")  # matches ending vowels
starting_vowel_match = 0
ending_vowel_match = 0

for line in sys.stdin:
    line = line.strip()  # removes leading and trailing whitespace
    words = line.lower().split()  # splits the line into words and converts to lowercase
    for word in words:
        if len(word) == 1:
            print(word[0], 1, *((1, 1) if word[0] in 'aeiou' else (0, 0))) # * unpacks startVowel 1 endVowel 1 if word[0] is a vowel
        else:
            print(word[0], 1, 1 if word[0] in 'aeiou' else 0, 0) 
            print(*(f'{letter} 1 0 0' for letter in word[1: -1]), sep='\n')
            print(word[-1], 1, 0, 1 if word[-1] in 'aeiou' else 0)

解决方案

核心是先过滤掉每个单词中的非字母字符,再基于纯字母序列执行原逻辑。修改后的代码如下:

import sys
import re

# 用集合存储元音,判断效率更高
VOWELS = {'a', 'e', 'i', 'o', 'u'}

for line in sys.stdin:
    line = line.strip()
    words = line.lower().split()
    for word in words:
        # 提取单词中所有小写字母,自动过滤非字母字符
        letters = re.findall(r'[a-z]', word)
        if not letters:
            continue  # 无字母的内容直接跳过
        
        char_count = len(letters)
        if char_count == 1:
            char = letters[0]
            is_vowel = char in VOWELS
            print(char, 1, 1 if is_vowel else 0, 1 if is_vowel else 0)
        else:
            # 处理首字符
            first_char = letters[0]
            print(first_char, 1, 1 if first_char in VOWELS else 0, 0)
            # 处理中间字符
            for char in letters[1:-1]:
                print(char, 1, 0, 0)
            # 处理尾字符
            last_char = letters[-1]
            print(last_char, 1, 0, 1 if last_char in VOWELS else 0)

关键修改点

  1. 过滤非字母字符:使用re.findall(r'[a-z]', word)提取单词内的所有小写字母(已提前转全小写),直接排除非字母符号。
  2. 空内容跳过:如果提取后的字母列表为空,直接跳过该单词,避免无效处理。
  3. 元音判断优化:用集合存储元音,比字符串in操作更快,逻辑更清晰。
  4. 基于纯字母序列处理:所有输出逻辑都基于过滤后的字母列表,确保输出仅包含字母。

测试效果

输入文本It's a beautiful life时:

  • "It's"会被提取为['i','t','s'],按规则输出这三个字符的对应标记
  • "a"输出a 1 1 1
  • "beautiful"提取全部字母后,依次输出每个字符的标记
  • "life"提取为['l','i','f','e'],输出对应标记

内容的提问来源于stack exchange,提问作者Zaku

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 14:45:35