You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Relationship列约束打乱Pandas的Text2列构建BERT负样本

构建带类别约束的BERT负样本方法

核心思路

生成负样本时,确保每个Text1对应的打乱后Text2,其Relationship类别与Text1完全不同;原样本标注为1,负样本标注为0。

具体实现步骤(基于Pandas)

  • 复制原数据集作为负样本基础,将标注统一设为0
  • 按Relationship列分组,建立「类别-对应Text2列表」的映射关系
  • 对每条样本,从除自身类别外的所有其他类别的Text2集合中随机抽取替换原Text2
  • 合并原正样本与处理后的负样本,得到最终训练集

代码示例

import pandas as pd
import numpy as np

# 示例原数据集
df = pd.DataFrame({
    'Text1': ['苹果树是落叶乔木', '胡萝卜是根茎类', '老虎是猫科动物'],
    'Text2': ['梨树也是落叶乔木', '白萝卜也是根茎类', '狮子也是猫科动物'],
    'Relationship': ['P', 'V', 'A'],
    'Label': [1, 1, 1]
})

# 构建类别到Text2列表的映射
rel_text2_map = df.groupby('Relationship')['Text2'].apply(list).to_dict()
all_rels = list(rel_text2_map.keys())

# 生成负样本
negative_df = df.copy()
negative_df['Label'] = 0

def get_neg_text2(row):
    current_rel = row['Relationship']
    # 收集所有非当前类别的Text2
    candidate_texts = []
    for rel in all_rels:
        if rel != current_rel:
            candidate_texts.extend(rel_text2_map[rel])
    # 随机选取一个
    return np.random.choice(candidate_texts)

negative_df['Text2'] = negative_df.apply(get_neg_text2, axis=1)

# 合并正负样本
final_df = pd.concat([df, negative_df], ignore_index=True)

优化提示

  • 数据集较大时,可改用向量化操作替代apply提升效率
  • 若某类别Text2数量过少,可设置最小候选数阈值,避免过度重复采样
  • 可多次执行打乱逻辑生成多组负样本,增强数据集丰富度

内容的提问来源于stack exchange,提问作者Chukwudi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 08:03:11