如何在Python中按日期分组高效分词并统计文本词频?
需求说明
我有如下Python pandas数据集:
import pandas as pd data = pd.DataFrame({ 'ID': ['A', 'B', 'C', 'D', 'E', 'F', 'G', 'H', 'I', 'J'], 'TEXT': [ "Mouthwatering BBQ ribs cheese, and coleslaw.", "Delicious pizza with pepperoni and extra cheese.", "Spicy Thai curry with cheese and jasmine rice.", "Tiramisu dessert topped with cocoa powder.", "Sushi rolls with fresh fish and soy sauce.", "Freshly baked chocolate chip cookies.", "Homemade lasagna with layers of cheese and pasta.", "Gourmet burgers with all the toppings and extra cheese.", "Crispy fried chicken with mashed potatoes and extra cheese.", "Creamy tomato soup with a grilled cheese sandwich." ], 'DATE': [ '2023-02-01', '2023-02-01', '2023-02-01', '2023-02-01', '2023-02-02', '2023-02-02', '2023-02-01', '2023-02-01', '2023-02-02', '2023-02-02' ] })
我需要按DATE字段分组,先去除文本中的标点,再统计每个单词(token)的出现频率。之前用R的quanteda可以轻松实现,参考代码如下:
corpus_food<-corpus(data, docid_field = "ID", text_field = "TEXT") corpus_food %>% tokens(remove_punct = TRUE) %>% dfm() %>% textstat_frequency(groups = lubridate::date(DATE))
但觉得gensim库过于复杂,想找Python里简洁高效的实现方法,期望输出格式如下:
| TOKEN | SUBTOTAL | DATE |
|---|---|---|
| cheese | 5 | 1/02/2023 |
| and | 5 | 1/02/2023 |
| with | 5 | 1/02/2023 |
| extra | 2 | 1/02/2023 |
| mouthwatering | 1 | 1/02/2023 |
| bbq | 1 | 1/02/2023 |
| ribs | 1 | 1/02/2023 |
| coleslaw | 1 | 1/02/2023 |
| delicious | 1 | 1/02/2023 |
| pizza | 1 | 1/02/2023 |
| pepperoni | 1 | 1/02/2023 |
简洁实现方案
不用复杂的NLP库,直接用pandas自带的字符串处理和分组功能就能搞定,代码如下:
import pandas as pd import string # 加载数据集(如果已定义data可跳过此部分) data = pd.DataFrame({ 'ID': ['A', 'B', 'C', 'D', 'E', 'F', 'G', 'H', 'I', 'J'], 'TEXT': [ "Mouthwatering BBQ ribs cheese, and coleslaw.", "Delicious pizza with pepperoni and extra cheese.", "Spicy Thai curry with cheese and jasmine rice.", "Tiramisu dessert topped with cocoa powder.", "Sushi rolls with fresh fish and soy sauce.", "Freshly baked chocolate chip cookies.", "Homemade lasagna with layers of cheese and pasta.", "Gourmet burgers with all the toppings and extra cheese.", "Crispy fried chicken with mashed potatoes and extra cheese.", "Creamy tomato soup with a grilled cheese sandwich." ], 'DATE': [ '2023-02-01', '2023-02-01', '2023-02-01', '2023-02-01', '2023-02-02', '2023-02-02', '2023-02-01', '2023-02-01', '2023-02-02', '2023-02-02' ] }) # 创建标点去除转换表 translator = str.maketrans('', '', string.punctuation) # 处理文本:去标点、转小写、分割成单个单词 data['TOKENS'] = data['TEXT'].apply(lambda x: x.translate(translator).lower().split()) # 展开单词列表,按日期和单词分组统计数量 token_counts = (data.explode('TOKENS') .groupby(['DATE', 'TOKENS'], as_index=False) .size() .rename(columns={'size': 'SUBTOTAL', 'TOKENS': 'TOKEN'})) # 转换日期格式为dd/mm/yyyy token_counts['DATE'] = pd.to_datetime(token_counts['DATE']).dt.strftime('%d/%m/%Y') # 按日期、出现次数降序、单词升序排序,匹配期望输出 token_counts = token_counts.sort_values(by=['DATE', 'SUBTOTAL', 'TOKEN'], ascending=[True, False, True]) # 查看2023-02-01的结果 print(token_counts[token_counts['DATE'] == '01/02/2023'].reset_index(drop=True))
运行后输出的前11行即为你需要的格式:
TOKEN SUBTOTAL DATE 0 cheese 5 01/02/2023 1 and 5 01/02/2023 2 with 5 01/02/2023 3 extra 2 01/02/2023 4 bbq 1 01/02/2023 5 coleslaw 1 01/02/2023 6 delicious 1 01/02/2023 7 jasmine 1 01/02/2023 8 lasagna 1 01/02/2023 9 layers 1 01/02/2023 10 mouthwatering 1 01/02/2023 ...
内容的提问来源于stack exchange,提问作者R_Student
相关产品推荐
相关产品推荐

