Python中如何用正则表达式提取"like"的各类变形单词?
提取指定单词的变形形式
需求
在给定的句子列表中搜索指定单词,提取该单词的各类变形形式(支持单次或多次出现)。
示例数据
输入的句子列表:
strings = ["He likes walking","I like that","He said he liked the movie"]
搜索关键词like,期望输出:
["likes","like","liked"]
尝试的代码及问题
之前写的代码:
keyword = "like" for string in strings: p = regex.search(keyword+".*" ,string,flags=regex.IGNORECASE) print(p.allcaptures()) print(p.group())
问题:这段代码会匹配like之后的所有内容,无法精准提取到like的各类变形单词。
解决方案
要精准匹配like的变形,需要用正则匹配完整的单词,确保只捕获以like开头的独立单词(忽略大小写)。可以用regex.findall批量提取每个句子中的目标单词,代码如下:
import regex strings = ["He likes walking","I like that","He said he liked the movie"] keyword = "like" results = [] for string in strings: # \b 匹配单词边界,避免匹配到包含like的非目标单词(比如likeness) # regex.escape(keyword) 转义关键词中的特殊字符,避免正则语法冲突 # \w* 匹配like之后的任意字母数字后缀,覆盖s、d这类常见变位后缀 matches = regex.findall(r'\b' + regex.escape(keyword) + r'\w*\b', string, flags=regex.IGNORECASE) results.extend(matches) print(results) # 输出: ['likes', 'like', 'liked']
内容的提问来源于stack exchange,提问作者A3006
相关产品推荐
相关产品推荐

