You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于日期、具体日期及主题语义匹配的双CSV文件合并方案咨询

可行的实现思路拆解

嘿,给你把问题拆解成三个核心步骤,每个步骤都附上具体思路和代码示例,应该能帮你走出困境:

一、先完成Specific_Date的范围缩小(含缺失值处理)

既然你已经按Date完成了分组,接下来可以在每个Date分组内,通过Specific_Date进一步过滤匹配范围,减少后续语义匹配的计算量:

  • 对于**双方都有有效Specific_Date**的行:把GPO和CAP中Specific_Date完全一致的行归为同一个子组,后续只需要在这个子组内做语义匹配。
  • 对于某一方或双方Specific_Date为NaN的行:把这些行单独放在Date分组的“缺失值子组”里,后续直接在整个Date分组内做语义匹配。

用Python pandas实现的参考代码:

import pandas as pd
import numpy as np

# 读取原始数据
gpo_df = pd.read_csv("gpo_full.csv")
cap_df = pd.read_csv("CAP_cols.csv")

# 按Date分组
gpo_groups = gpo_df.groupby("Date")
cap_groups = cap_df.groupby("Date")

# 存储分组后的匹配候选池
match_candidates = {}

for date, gpo_group in gpo_groups:
    if date not in cap_groups.groups:
        continue  # CAP中无该Date数据,直接跳过
    cap_group = cap_groups.get_group(date)
    
    # 1. 处理有有效Specific_Date的子组
    gpo_valid = gpo_group[gpo_group["Specific_Date"].notna()]
    cap_valid = cap_group[cap_group["Specific_Date"].notna()]
    
    for spec_date, gpo_sub in gpo_valid.groupby("Specific_Date"):
        cap_sub = cap_valid[cap_valid["Specific_Date"] == spec_date]
        if not cap_sub.empty:
            key = (date, spec_date)
            match_candidates[key] = {"gpo": gpo_sub, "cap": cap_sub}
    
    # 2. 处理Specific_Date有缺失的子组
    gpo_missing = gpo_group[gpo_group["Specific_Date"].isna()]
    if not gpo_missing.empty:
        key = (date, np.nan)
        # 缺失值的GPO行需要在整个Date分组的CAP数据中找匹配
        match_candidates[key] = {"gpo": gpo_missing, "cap": cap_group}

二、基于词嵌入的主题语义匹配

这一步核心是给每个topic生成语义向量,通过相似度计算为每个GPO行匹配最合适的CAP行。这里推荐用Sentence-BERT——它比传统Word2Vec更适合短文本语义匹配,预训练模型直接可用,上手成本低:

实现步骤:

  1. 先安装依赖:pip install sentence-transformers
  2. 加载轻量高效的预训练模型(比如all-MiniLM-L6-v2)
  3. 为每个候选子组的GPO、CAPtopic生成语义向量
  4. 计算余弦相似度,为每个GPO行挑选相似度最高的CAP行(若要求CAP行唯一匹配,可改用匈牙利算法做一对一最优分配)

参考代码:

from sentence_transformers import SentenceTransformer, util
from scipy.optimize import linear_sum_assignment

# 加载预训练语义匹配模型
model = SentenceTransformer('all-MiniLM-L6-v2')

# 存储最终匹配结果
final_results = []

for (date, spec_date), candidates in match_candidates.items():
    gpo_sub = candidates["gpo"]
    cap_sub = candidates["cap"]
    
    # 处理topic的NaN值,替换为空字符串避免模型报错
    gpo_topics = gpo_sub["topic"].fillna("").tolist()
    cap_topics = cap_sub["topic"].fillna("").tolist()
    
    # 生成语义向量
    gpo_embeddings = model.encode(gpo_topics, convert_to_tensor=True)
    cap_embeddings = model.encode(cap_topics, convert_to_tensor=True)
    
    # 计算余弦相似度矩阵
    cos_scores = util.cos_sim(gpo_embeddings, cap_embeddings).cpu().numpy()
    
    # 方案1:一对一最优匹配(保证每个CAP行只被匹配一次)
    cost_matrix = 1 - cos_scores  # 转成成本矩阵,用匈牙利算法找最小成本
    row_ind, col_ind = linear_sum_assignment(cost_matrix)
    for gpo_idx, cap_idx in zip(row_ind, col_ind):
        gpo_row = gpo_sub.iloc[gpo_idx]
        cap_row = cap_sub.iloc[cap_idx]
        
        # 构建目标格式的结果项
        result_item = {
            "Date": date,
            "Specific_Date": spec_date if not pd.isna(spec_date) else None,
            "Topic": {
                "GPO": gpo_row["topic"] if pd.notna(gpo_row["topic"]) else None,
                "CAP": cap_row["topic"] if pd.notna(cap_row["topic"]) else None
            },
            "GPO": {
                "hearing_sub_type": gpo_row["hearing_sub_type"],
                # 可添加更多GPO字段
            },
            "CAP": {
                "majortopic": cap_row["majortopic"],
                "id": cap_row["id"],
                "Chamber": cap_row["Chamber"],
                # 可添加更多CAP字段
            }
        }
        final_results.append(result_item)
    
    # 方案2:每个GPO取相似度最高的CAP行(允许CAP行被重复匹配)
    # for gpo_idx, gpo_row in gpo_sub.iterrows():
    #     top_cap_idx = cos_scores[gpo_idx].argmax()
    #     cap_row = cap_sub.iloc[top_cap_idx]
    #     # 构建result_item并添加到final_results

三、导出为目标JSON格式

最后把结果导出成你需要的JSON格式即可:

import json

with open("matched_result.json", "w", encoding="utf-8") as f:
    json.dump(final_results, f, indent=2, default=str)

内容的提问来源于stack exchange,提问作者whitneyl

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 22:17:52