如何用Python apply-map实现Pandas DataFrame段落的最长单词统计?
解决Pandas文本列提取最长单词信息的问题
我来帮你搞定这个需求!我们可以通过自定义一个处理函数,结合apply方法来生成你需要的三列数据。关键是先把每个段落里的单词清洗干净,再统计最长单词的相关信息。
步骤详解
1. 导入必要的库
首先要导入pandas和处理文本需要的string模块(用来处理标点):
import pandas as pd import string
2. 定义处理单个文本的函数
这个函数会接收一段文本,返回包含最长单词长度、最长单词数量、所有最长单词的元组:
def process_text(text): # 分割文本为单词,同时去掉每个单词前后的标点 words = [word.strip(string.punctuation) for word in text.split()] # 过滤掉分割后可能出现的空字符串(比如连续标点的情况) words = [word for word in words if word] if not words: # 处理空文本的边界情况,避免报错 return (0, 0, "") # 计算每个单词的长度 word_lengths = [len(word) for word in words] max_length = max(word_lengths) # 筛选出所有长度等于max_length的单词 longest_words = [word for word, length in zip(words, word_lengths) if length == max_length] return (max_length, len(longest_words), " ".join(longest_words))
3. 应用函数到DataFrame并生成新列
用apply方法把函数应用到text列,然后把返回的元组拆分成三列:
# 你的示例DataFrame df = pd.DataFrame({'text':[ "that's not where the biggest opportunity is - it's with heart failure drug - very very huge market....", "Of course! I just got diagnosed with congestive heart failure and type 2 diabetes. I smoked for 12 years and ate like crap for about the same time. I quit smoking and have been on a diet for a few weeks now. Let me assure you that I'd rather have a coke, gummi bears, and a bag of cheez doodles than a pack of cigs right now. Addiction is addiction.", "STILLWATER, Okla. (AP) ? Medical examiner spokeswoman SpokesWoman: Oklahoma State player Tyrek Coger died of enlarged heart, manner of death ruled natural." ]}) # 应用函数并拆分列 df[['word_length', 'word_count', 'words']] = df['text'].apply(lambda x: pd.Series(process_text(x))) # 调整列顺序(和你的预期输出结构一致) df = df[['text', 'word_count', 'word_length', 'words']]
4. 查看最终结果
运行后你会得到和预期完全匹配的输出:
text word_count word_length words 0 that's not where the biggest opportunity is - ... 1 11 opportunity 1 Of course! I just got diagnosed with congestiv... 1 10 congestive 2 STILLWATER, Okla. (AP) ? Medical examiner spok... 2 11 spokeswoman SpokesWoman
关键细节说明
- 标点处理:用
word.strip(string.punctuation)去掉每个单词前后的标点,既保留了that's这类带撇号的单词,又能清理掉opportunity....、STILLWATER,这类多余符号。 - 边界情况:函数里处理了空文本的特殊场景,避免出现
max()函数报错的问题。 - 大小写区分:示例第三行的
spokeswoman和SpokesWoman被视为两个独立单词,完全符合你的预期输出要求。
内容的提问来源于stack exchange,提问作者sync11
相关产品推荐
相关产品推荐

