如何在Python的DataFrame中提取以‘no’开头的目标短语?
问题描述
已从DataFrame中筛选出包含单词"no"的列,需要提取并打印其中以"no"开头的短语。示例数据集如下:
index | Column 1 ------------------------------------------------------------------------ 0 | no school for the rest of the year. no homework and no classes 1 | no more worries. no stress and no more anxiety 2 | no teachers telling us what to do
期望输出:
no school no homework no classes no more worries no stress no more anxiety no teachers
当前代码:
# 复制目标列 copy = df4['phrases'].copy() # 筛选含"no"的行 nomore = copy.str.contains(r'\bno\b', na=False) # 拆分每行字符串为单词列表 copy.loc[nomore] = copy[nomore].str.split()
尝试以下代码拼接短语失败:
for i in copy.loc[nomore]: for x in i: if x == 'no': print(x,x+1)
问题:无法正确匹配x == 'no',且x+1报错,求修复方法。
解决方案
原代码问题分析
- 标点干扰:拆分后的单词可能附带标点(如
school.),不过x == 'no'本身逻辑没问题,但x是字符串,x+1属于错误的字符串拼接操作,并非取列表下一个元素。 - 短语长度不固定:有的短语是
no+单词结构,有的是no+多词结构(如no more worries),遍历单个单词的方式无法完整提取多词短语。
推荐方法:正则直接提取
不用拆分单词,用正则匹配所有符合要求的短语,高效且简洁:
import re import string # 遍历目标列的每一行文本 for text in df4['phrases']: # 匹配所有以独立单词no开头的短语,支持多词组合 phrases = re.findall(r'\bno\s+(?:\w+\s*)+', text) # 清理短语的多余空格和末尾标点后打印 for p in phrases: clean_phrase = p.strip().rstrip(string.punctuation) print(clean_phrase)
正则说明:
\bno:匹配独立的单词no,避免误匹配now这类包含no的单词\s+(?:\w+\s*)+:匹配no后面的一个或多个单词,允许单词间的空格rstrip(string.punctuation):去掉短语末尾的标点符号(如.)
备选:修复循环逻辑
如果坚持用拆分单词的方式,可调整代码如下:
import string copy = df4['phrases'].copy() nomore = copy.str.contains(r'\bno\b', na=False) copy.loc[nomore] = copy[nomore].str.split() for word_list in copy.loc[nomore]: idx = 0 while idx < len(word_list): # 清理单词的标点并转小写,判断是否为no word_clean = word_list[idx].strip(string.punctuation).lower() if word_clean == 'no': # 收集短语内容,直到遇到下一个no或列表结束 phrase_words = [] current_idx = idx while current_idx < len(word_list): current_word = word_list[current_idx].strip(string.punctuation) phrase_words.append(current_word) current_idx += 1 # 提前终止:如果下一个单词是no,停止当前短语收集 if current_idx < len(word_list) and word_list[current_idx].strip(string.punctuation).lower() == 'no': break print(' '.join(phrase_words)) idx = current_idx # 跳到下一个no的位置,避免重复处理 else: idx += 1
这段代码通过索引遍历解决了标点干扰问题,同时能完整提取多词短语。
内容的提问来源于stack exchange,提问作者shorttriptomars
相关产品推荐
相关产品推荐

