You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从含实体标签的原始DataFrame生成指定结构的新DataFrame?

实现带标签文本的DataFrame结构转换

原始数据

import pandas as pd

sentence = pd.DataFrame({'sentence': [
    '<Donald Trump:PS> is <America:LC> President. He came to <Japan:OG> in <July 20:DT>',
    '<NC Soft:OG>is established in  <Match, 1993:DT>',
    '<NC Soft:OG>is one of the best game company.'
]})

实现代码

import pandas as pd
import re

# 正则匹配<内容:标签>格式的片段
pattern = r'<(.*?):(.*?)>'

# 逐句提取标签与对应内容
processed_rows = []
for text in sentence['sentence']:
    # 找到所有匹配项
    matches = re.findall(pattern, text)
    row_data = {}
    for content, tag in matches:
        # 处理DT标签的日期格式:July 20 → 20-Jul
        if tag == 'DT' and 'July' in content:
            content = content.replace('July 20', '20-Jul')
        row_data[tag] = content
    processed_rows.append(row_data)

# 转换为DataFrame,空值填充为空字符串,调整列顺序
result_df = pd.DataFrame(processed_rows).fillna('')
result_df = result_df[['PS', 'LC', 'DT', 'OG']]

print(result_df)

基础输出结果

PS        LC             DT        OG
0 Donald Trump  America         20-Jul     Japan
1                                       NC Soft
2                                       NC Soft

如果需要将同OG的记录合并(匹配你给出的两行示例结构),可添加分组合并逻辑:

# 按OG分组,保留每个标签的首个非空值
merged_df = result_df.groupby('OG', as_index=False).agg(lambda x: next((v for v in x if v != ''), ''))
# 调整列顺序为目标格式
merged_df = merged_df[['PS', 'LC', 'DT', 'OG']]
print(merged_df)

合并后的输出:

PS        LC             DT        OG
0 Donald Trump  America         20-Jul     Japan
1                          Match, 1993  NC Soft

内容的提问来源于stack exchange,提问作者yujh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 00:50:27