DataFrame中Bigrams排序、去重及频次汇总的实现求助
Pandas DataFrame二元组去重并累加计数解决方案
问题场景
现有一个包含bigrams(字符串格式的二元组)和counts列的DataFrame,数据如下:
| bigrams | counts |
|---|---|
| ('asset', 'experience') | 1 |
| ('qualifications', 'your') | 1 |
| ('your', 'contribution') | 1 |
| ('contribution', 'bilingual') | 1 |
| ('your', 'qualifications') | 1 |
| ('bilingual', 'contribution') | 1 |
需求:
- 将
bigrams中的二元组按字母顺序排序(如('contribution', 'bilingual')转为('bilingual', 'contribution')) - 去除重复的二元组并累加对应
counts值 - 最终按
counts降序排列,同时保留bigrams原有字符串格式
正确解决方案代码
import pandas as pd # 读取数据 df = pd.read_csv('emplois_df_FonctionsStagiaire_bigrams_counts.csv') def process_bigram(bigram_str): # 去除首尾括号,分割出两个带引号的单词 words = bigram_str.strip('()').split(', ') # 清理单引号后排序,再重新组装成原格式字符串 sorted_words = sorted([word.strip("'") for word in words]) return f"('{sorted_words[0]}', '{sorted_words[1]}')" # 生成标准化排序后的二元组列 df['processed_bigram'] = df['bigrams'].apply(process_bigram) # 分组累加计数,再按counts降序排序 df_new = df.groupby('processed_bigram', as_index=False)['counts'].sum() df_new = df_new.sort_values(by='counts', ascending=False) # 恢复原列名 df_new.rename(columns={'processed_bigram': 'bigrams'}, inplace=True) # 保存结果 df_new.to_csv('emplois_df_FonctionsStagiaire_bigrams_sorted_counts.csv', index=False)
代码说明
- process_bigram函数:针对性处理字符串格式的二元组,先清理括号、单引号等冗余符号,提取纯单词后排序,再还原成原有的
('word1', 'word2')格式,确保分组时重复项能被正确识别。 - 分组累加:以标准化后的二元组为分组依据,对
counts求和,完成去重累加操作。 - 排序输出:最后按
counts降序排列,匹配需求中的输出顺序。
内容的提问来源于stack exchange,提问作者gmohor21
相关产品推荐
相关产品推荐

