You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何拆分含'|'分隔符的特殊格式字符串并转化为指定结构的DataFrame

实现思路

  • 统一处理所有输入字符串,拆分得到所有「词对-关系」三元组
  • 按word1和word2对三元组分组,收集同一词对的所有关系
  • 每个词对的关系列表补空值到固定长度5,最后拼接为目标DataFrame

完整实现代码

import pandas as pd
import numpy as np

# 定义输入字符串
str1 = 'ABC ; BC /contain/lower | ABC ; BC /notsame | ABC ; BC /similar\n'
str2 = 'AB ; FC /notsame\n'
input_strs = [str1, str2]

# 存储所有三元组的列表
triples = []

for s in input_strs:
    # 去除换行符后按|拆分所有条目,清除每个条目的首尾冗余空格
    entries = [entry.strip() for entry in s.strip().split('|')]
    for entry in entries:
        # 拆分提取word1
        part1, part2 = entry.split(';', maxsplit=1)
        word1 = part1.strip()
        # 拆分提取word2和relation
        word2, relation = part2.strip().split(' ', maxsplit=1)
        triples.append((word1, word2, relation))

# 按词对分组收集所有关系
grouped = {}
for w1, w2, rel in triples:
    key = (w1, w2)
    if key not in grouped:
        grouped[key] = []
    grouped[key].append(rel)

# 构建目标DataFrame
rows = []
for (w1, w2), rels in grouped.items():
    # 关系列表补NaN到固定长度5
    padded_rels = rels + [np.nan] * (5 - len(rels))
    rows.append([w1, w2] + padded_rels)

df = pd.DataFrame(rows, columns=['word1', 'word2', 'relation1', 'relation2', 'relation3', 'relation4', 'relation5'])

运行后输出的df结构和你要求的完全一致。

内容的提问来源于stack exchange,提问作者Mark J.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 01:36:07