You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

DataFrame中Bigrams排序、去重及频次汇总的实现求助

Pandas DataFrame二元组去重并累加计数解决方案

问题场景

现有一个包含bigrams(字符串格式的二元组)和counts列的DataFrame,数据如下:

bigramscounts
('asset', 'experience')1
('qualifications', 'your')1
('your', 'contribution')1
('contribution', 'bilingual')1
('your', 'qualifications')1
('bilingual', 'contribution')1

需求:

  • 将bigrams中的二元组按字母顺序排序(如('contribution', 'bilingual')转为('bilingual', 'contribution'))
  • 去除重复的二元组并累加对应counts值
  • 最终按counts降序排列,同时保留bigrams原有字符串格式

正确解决方案代码

import pandas as pd

# 读取数据
df = pd.read_csv('emplois_df_FonctionsStagiaire_bigrams_counts.csv')

def process_bigram(bigram_str):
    # 去除首尾括号,分割出两个带引号的单词
    words = bigram_str.strip('()').split(', ')
    # 清理单引号后排序,再重新组装成原格式字符串
    sorted_words = sorted([word.strip("'") for word in words])
    return f"('{sorted_words[0]}', '{sorted_words[1]}')"

# 生成标准化排序后的二元组列
df['processed_bigram'] = df['bigrams'].apply(process_bigram)

# 分组累加计数,再按counts降序排序
df_new = df.groupby('processed_bigram', as_index=False)['counts'].sum()
df_new = df_new.sort_values(by='counts', ascending=False)

# 恢复原列名
df_new.rename(columns={'processed_bigram': 'bigrams'}, inplace=True)

# 保存结果
df_new.to_csv('emplois_df_FonctionsStagiaire_bigrams_sorted_counts.csv', index=False)

代码说明

  • process_bigram函数:针对性处理字符串格式的二元组,先清理括号、单引号等冗余符号,提取纯单词后排序,再还原成原有的('word1', 'word2')格式,确保分组时重复项能被正确识别。
  • 分组累加:以标准化后的二元组为分组依据,对counts求和,完成去重累加操作。
  • 排序输出:最后按counts降序排列,匹配需求中的输出顺序。

内容的提问来源于stack exchange,提问作者gmohor21

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 16:34:55