获取字符串指定字符位置前的n个完整单词(不截断单词)
获取指定字符位置前的n个完整单词
原代码的问题在于直接截取到char_nr-1时,可能会截断到某个单词中间,导致split后出现不完整单词。要解决这个问题,核心是确保我们只截取到完整单词的末尾,再进行分割。
方法一:利用空格定位(简单高效)
通过找到目标字符位置前的最后一个空格,截取到该位置的子串,保证所有内容都是完整单词:
text = 'the house is big the house is big the house is big' char_nr = 19 nr_words = 3 # 从0到char_nr范围内,反向查找最后一个空格的索引 last_space_pos = text.rfind(' ', 0, char_nr) # 截取到该空格位置,得到仅含完整单词的字符串 full_words_text = text[:last_space_pos] # 分割为单词列表 word_list = full_words_text.split() # 处理n大于单词总数的情况 if nr_words > len(word_list): nr_words = len(word_list) result = word_list[-nr_words:] print(result) # 输出: ['house', 'is', 'big']
方法二:正则表达式匹配(更灵活)
如果需要处理复杂的单词边界(比如包含连字符、缩写符号的单词),可以用正则匹配所有完整单词的位置,筛选出结束位置在char_nr之前的单词:
import re text = 'the house is big the house is big the house is big' char_nr = 19 nr_words = 3 # 匹配所有单词,获取每个单词的起止位置 word_matches = list(re.finditer(r'\b\w+\b', text)) # 筛选出结束位置小于char_nr的完整单词 full_words = [match.group() for match in word_matches if match.end() < char_nr] # 取最后n个单词 if nr_words > len(full_words): nr_words = len(full_words) result = full_words[-nr_words:] print(result) # 输出: ['house', 'is', 'big']
两种方法都能解决你的问题,方法一更适合普通空格分隔的文本,方法二适配更复杂的单词场景。
内容的提问来源于stack exchange,提问作者JFerro
相关产品推荐
相关产品推荐

