如何遍历文档列表提取所有以指定关键词开头的单词?
解决方法
原代码的问题在于re.search仅会返回每行中的第一个匹配结果,因此会漏掉同一行内多次出现的目标单词。要获取所有匹配项,需要改用re.findall,同时调整正则表达式以适配单词在句首、句尾的情况:
import re doc_list = ['welcome to python tutorial', 'you are welcome to python_one tutorial', 'welcome python_two tutorial python_three tutorial', 'hello python', 'hello python_four world'] output = [] keyword = 'python' for line in doc_list: # 匹配所有以指定关键词开头的完整单词 matches = re.findall(r'\b' + re.escape(keyword) + r'\w*\b', line) output.extend(matches) print(output)
关键改进点:
- 用
re.findall替代re.search:该方法会返回当前字符串中所有符合正则规则的匹配项,而非仅第一个 - 使用
\b单词边界:确保匹配的是完整单词,避免误匹配包含关键词的其他字符串(比如xpython) re.escape(keyword):对关键词进行转义,防止关键词中包含正则特殊字符(如.、*)时导致匹配出错\w*:匹配关键词后续的字母、数字或下划线,适配python_one这类带后缀的单词格式
运行上述代码后,输出结果为:['python', 'python_one', 'python_two', 'python_three', 'python', 'python_four'],完全覆盖所有目标单词。
内容的提问来源于stack exchange,提问作者Milind
相关产品推荐
相关产品推荐

