求助:字符串首尾标点去除实现及单词检测函数测试异常
问题求助:字符串首尾标点处理及单词检测函数异常
需求一:实现去除字符串首尾标点的函数
需要编写一个函数,能够移除字符串首尾的标点符号。示例:输入 "_cat's_",处理后得到 "cat's"。
需求二:修复单词检测函数的异常
我写了一个detect_word函数,用于检测指定单词是否存在于源字符串中。最初没使用strip()方法,查资料后添加了该方法,但现在第一个测试用例始终返回False,不符合预期(预期前两个测试用例返回True,第三个返回False)。
现有代码
import string def detect_word(source, word): source = source.strip().lower() word = word.strip().lower() split_source = source.split() if word in split_source: return True else: return False
测试用例
t1 = "I have a cat." t2 = "My cat is orange" t3 = "Everything is a catastrophe." print(detect_word(t1, 'cat')) # 预期True,实际False print(detect_word(t2, 'cat')) # 预期True,实际True print(detect_word(t3, 'cat')) # 预期False,实际False
问题原因
问题出在strip()的默认行为上:默认的strip()只会移除字符串首尾的空白字符(空格、换行、制表符等),不会处理单词末尾或中间的标点。比如测试用例t1处理后是"i have a cat.",split()得到的列表里是["i", "have", "a", "cat."],和目标单词"cat"不匹配,因此返回False。
解决方案
1. 修复detect_word函数
要实现准确的单词匹配,需要移除每个单词首尾的标点后再做比较,而不是仅处理整个字符串的首尾:
import string def detect_word(source, word): # 处理目标单词:移除首尾标点并转小写 target_word = word.strip(string.punctuation).lower() # 遍历源字符串分割后的每个单词,处理后对比 for token in source.lower().split(): cleaned_token = token.strip(string.punctuation) if cleaned_token == target_word: return True return False
2. 实现去除字符串首尾标点的函数
利用string.punctuation判断字符是否为标点,循环移除首尾的标点符号:
import string def strip_punctuation(s): # 从开头移除标点 start = 0 while start < len(s) and s[start] in string.punctuation: start += 1 # 从结尾移除标点 end = len(s) - 1 while end >= start and s[end] in string.punctuation: end -= 1 return s[start:end+1] # 测试示例 print(strip_punctuation("_cat's_")) # 输出: "cat's"
修复后测试结果
运行修复后的detect_word函数,测试用例结果将符合预期:
detect_word(t1, 'cat')返回Truedetect_word(t2, 'cat')返回Truedetect_word(t3, 'cat')返回False
内容的提问来源于stack exchange,提问作者Sukwha
相关产品推荐
相关产品推荐

