You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何统计数据集中两种农药同日期同地点共现次数并按频次排序?

解决方案(基于R语言tidyverse)

假设你的数据集名为df,以下是实现需求的分步代码和思路:

步骤1:数据预处理(按需选择)

如果同一日期、地点下的同一农药多条记录仅需标记“存在”而非重复计数,先对数据去重:

library(tidyverse)

# 保留每个(日期+地点+农药)的唯一记录
unique_df <- df %>%
  distinct(date, monitoring_location_id, bma_compound_name)

如果需要统计同一日期地点下农药记录数的乘积总和(比如A有3条、B有2条,该组贡献6次共现),则直接统计每个(日期+地点+农药)的记录数:

count_df <- df %>%
  count(date, monitoring_location_id, bma_compound_name, name = "record_num")

步骤2:生成无序农药组合

核心是让A-B和B-A视为同一组合,通过sort()对农药名称排序后再生成组合,确保组合表述统一:

场景1:统计共现的(日期+地点)事件数

# 按日期+地点分组,生成每组内的两两农药组合
pair_events <- unique_df %>%
  group_by(date, monitoring_location_id) %>%
  filter(n() >= 2)  # 过滤仅含单一农药的组
  summarise(pair = list(combn(sort(bma_compound_name), 2, paste, collapse = "-")), .groups = "drop") %>%
  unnest(pair)

场景2:统计记录数乘积的总和

pair_products <- count_df %>%
  group_by(date, monitoring_location_id) %>%
  filter(n() >= 2) %>%
  summarise(
    pair = list(combn(sort(bma_compound_name), 2, paste, collapse = "-")),
    product = list(combn(record_num, 2, prod)),
    .groups = "drop"
  ) %>%
  unnest(c(pair, product))

步骤3:统计频次并排序

场景1:统计事件数

result_events <- pair_events %>%
  count(pair, name = "cooccurrence_count") %>%
  arrange(desc(cooccurrence_count))

场景2:统计乘积总和

result_products <- pair_products %>%
  group_by(pair) %>%
  summarise(total_cooccurrence = sum(product), .groups = "drop") %>%
  arrange(desc(total_cooccurrence))

关键说明

  • sort()处理农药名称后生成组合,彻底避免了正反组合重复统计的问题;
  • filter(n() >= 2)过滤无效组,减少不必要计算;
  • 最终结果自动排除从未共现的组合,只有实际共现的组合会被统计。

内容的提问来源于stack exchange,提问作者ramateur

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 06:04:56