You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python apply-map实现Pandas DataFrame段落的最长单词统计?

解决Pandas文本列提取最长单词信息的问题

我来帮你搞定这个需求!我们可以通过自定义一个处理函数,结合apply方法来生成你需要的三列数据。关键是先把每个段落里的单词清洗干净,再统计最长单词的相关信息。

步骤详解

1. 导入必要的库

首先要导入pandas和处理文本需要的string模块(用来处理标点):

import pandas as pd
import string

2. 定义处理单个文本的函数

这个函数会接收一段文本,返回包含最长单词长度、最长单词数量、所有最长单词的元组:

def process_text(text):
    # 分割文本为单词,同时去掉每个单词前后的标点
    words = [word.strip(string.punctuation) for word in text.split()]
    # 过滤掉分割后可能出现的空字符串(比如连续标点的情况)
    words = [word for word in words if word]
    
    if not words:  # 处理空文本的边界情况,避免报错
        return (0, 0, "")
    
    # 计算每个单词的长度
    word_lengths = [len(word) for word in words]
    max_length = max(word_lengths)
    
    # 筛选出所有长度等于max_length的单词
    longest_words = [word for word, length in zip(words, word_lengths) if length == max_length]
    
    return (max_length, len(longest_words), " ".join(longest_words))

3. 应用函数到DataFrame并生成新列

用apply方法把函数应用到text列,然后把返回的元组拆分成三列:

# 你的示例DataFrame
df = pd.DataFrame({'text':[ "that's not where the biggest opportunity is - it's with heart failure drug - very very huge market....", "Of course! I just got diagnosed with congestive heart failure and type 2 diabetes. I smoked for 12 years and ate like crap for about the same time. I quit smoking and have been on a diet for a few weeks now. Let me assure you that I'd rather have a coke, gummi bears, and a bag of cheez doodles than a pack of cigs right now. Addiction is addiction.", "STILLWATER, Okla. (AP) ? Medical examiner spokeswoman SpokesWoman: Oklahoma State player Tyrek Coger died of enlarged heart, manner of death ruled natural." ]})

# 应用函数并拆分列
df[['word_length', 'word_count', 'words']] = df['text'].apply(lambda x: pd.Series(process_text(x)))

# 调整列顺序(和你的预期输出结构一致)
df = df[['text', 'word_count', 'word_length', 'words']]

4. 查看最终结果

运行后你会得到和预期完全匹配的输出:

text  word_count  word_length               words
0  that's not where the biggest opportunity is - ...           1           11        opportunity
1  Of course! I just got diagnosed with congestiv...           1           10        congestive
2  STILLWATER, Okla. (AP) ? Medical examiner spok...           2           11  spokeswoman SpokesWoman

关键细节说明

  • 标点处理:用word.strip(string.punctuation)去掉每个单词前后的标点,既保留了that's这类带撇号的单词,又能清理掉opportunity....、STILLWATER,这类多余符号。
  • 边界情况:函数里处理了空文本的特殊场景,避免出现max()函数报错的问题。
  • 大小写区分:示例第三行的spokeswoman和SpokesWoman被视为两个独立单词,完全符合你的预期输出要求。

内容的提问来源于stack exchange,提问作者sync11

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 09:13:04