You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python的DataFrame中提取以‘no’开头的目标短语?

问题描述

已从DataFrame中筛选出包含单词"no"的列,需要提取并打印其中以"no"开头的短语。示例数据集如下:

index | Column 1
------------------------------------------------------------------------ 
  0   | no school for the rest of the year. no homework and no classes
  1   | no more worries. no stress and no more anxiety
  2   | no teachers telling us what to do

期望输出:

no school
no homework
no classes
no more worries
no stress
no more anxiety
no teachers

当前代码:

# 复制目标列
copy = df4['phrases'].copy()

# 筛选含"no"的行
nomore = copy.str.contains(r'\bno\b', na=False)

# 拆分每行字符串为单词列表
copy.loc[nomore] = copy[nomore].str.split()

尝试以下代码拼接短语失败:

for i in  copy.loc[nomore]:
    for x in i: 
        if x == 'no':
            print(x,x+1)

问题:无法正确匹配x == 'no',且x+1报错,求修复方法。

解决方案

原代码问题分析

  1. 标点干扰:拆分后的单词可能附带标点(如school.),不过x == 'no'本身逻辑没问题,但x是字符串,x+1属于错误的字符串拼接操作,并非取列表下一个元素。
  2. 短语长度不固定:有的短语是no+单词结构,有的是no+多词结构(如no more worries),遍历单个单词的方式无法完整提取多词短语。

推荐方法:正则直接提取

不用拆分单词,用正则匹配所有符合要求的短语,高效且简洁:

import re
import string

# 遍历目标列的每一行文本
for text in df4['phrases']:
    # 匹配所有以独立单词no开头的短语,支持多词组合
    phrases = re.findall(r'\bno\s+(?:\w+\s*)+', text)
    # 清理短语的多余空格和末尾标点后打印
    for p in phrases:
        clean_phrase = p.strip().rstrip(string.punctuation)
        print(clean_phrase)

正则说明:

  • \bno:匹配独立的单词no,避免误匹配now这类包含no的单词
  • \s+(?:\w+\s*)+:匹配no后面的一个或多个单词,允许单词间的空格
  • rstrip(string.punctuation):去掉短语末尾的标点符号(如.)

备选:修复循环逻辑

如果坚持用拆分单词的方式,可调整代码如下:

import string

copy = df4['phrases'].copy()
nomore = copy.str.contains(r'\bno\b', na=False)
copy.loc[nomore] = copy[nomore].str.split()

for word_list in copy.loc[nomore]:
    idx = 0
    while idx < len(word_list):
        # 清理单词的标点并转小写,判断是否为no
        word_clean = word_list[idx].strip(string.punctuation).lower()
        if word_clean == 'no':
            # 收集短语内容,直到遇到下一个no或列表结束
            phrase_words = []
            current_idx = idx
            while current_idx < len(word_list):
                current_word = word_list[current_idx].strip(string.punctuation)
                phrase_words.append(current_word)
                current_idx += 1
                # 提前终止:如果下一个单词是no,停止当前短语收集
                if current_idx < len(word_list) and word_list[current_idx].strip(string.punctuation).lower() == 'no':
                    break
            print(' '.join(phrase_words))
            idx = current_idx  # 跳到下一个no的位置,避免重复处理
        else:
            idx += 1

这段代码通过索引遍历解决了标点干扰问题,同时能完整提取多词短语。

内容的提问来源于stack exchange,提问作者shorttriptomars

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 22:53:15