You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将非年度政党意识形态数据匹配到年度执政数据框?

基于时间推断的数据匹配合并方案

问题背景

现有两个数据框:

  • df1:记录各国每一年的执政政党信息(例如美国2017-2021年共和党执政)
  • df2:记录政党意识形态的时间变更情况,但并非逐年记录(例如政党A1970年为左翼,1980年变为中左翼),二者政党编码一致。

直接使用dplyr::left_join()合并会出现大量NA值,因为df2没有逐年数据。需要根据df2的时间节点,推断出df1对应年份的意识形态。

示例数据

df1示例

Country | Year | Government's Political Party ID
X       | 1990 | 340
X       | 1991 | 340
X       | 1992 | 340
X       | 1993 | 340

df2示例

Country | Year | Political Party ID | Ideology
X       | 1970 | 340                | center
X       | 1985 | 340                | center
X       | 1992 | 340                | center-left
X       | 1999 | 340                | center-left

期望合并结果

Country | Year | Government's Political Party ID | Ideology
X       | 1990 | 340                             | center
X       | 1991 | 340                             | center-left
X       | 1992 | 340                             | center-left
X       | 1993 | 340                             | center-left

解决方案

核心逻辑是:按国家+政党ID分组,为df1的每个年份匹配df2中小于等于该年份的最近时间点对应的意识形态。以下是两种实用实现方法:

方法1:使用fuzzyjoin包的模糊匹配

fuzzyjoin支持基于条件的模糊匹配,能高效实现时间区间的匹配:

  1. 加载依赖包:
library(dplyr)
library(fuzzyjoin)
  1. 执行匹配与结果清理:
merged_df <- fuzzy_left_join(
  df1,
  df2,
  by = c(
    "Country" = "Country",
    "Government's Political Party ID" = "Political Party ID",
    "Year" = "Year"
  ),
  match_fun = list(`==`, `==`, `>=`)
) %>%
  group_by(Country, Year.x, `Government's Political Party ID`) %>%
  filter(Year.y == max(Year.y)) %>%
  ungroup() %>%
  rename(Year = Year.x) %>%
  select(Country, Year, `Government's Political Party ID`, Ideology)

方法2:仅使用dplyr实现

无需额外安装包,通过分组窗口函数完成匹配:

  1. 加载dplyr:
library(dplyr)
  1. 数据整理与匹配:
# 先对df2按国家、政党ID分组并按年份排序
df2_cleaned <- df2 %>%
  group_by(Country, `Political Party ID`) %>%
  arrange(Year) %>%
  ungroup()

# 合并后筛选每个df1行对应的最新df2记录
merged_df <- df1 %>%
  left_join(df2_cleaned, 
            by = c("Country" = "Country", 
                   "Government's Political Party ID" = "Political Party ID")) %>%
  group_by(Country, Year.x, `Government's Political Party ID`) %>%
  filter(Year.y <= Year.x) %>%
  filter(Year.y == max(Year.y)) %>%
  ungroup() %>%
  rename(Year = Year.x) %>%
  select(Country, Year, `Government's Political Party ID`, Ideology)

内容的提问来源于stack exchange,提问作者Pedro Cardoso

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 18:52:37