如何从含实体标签的原始DataFrame生成指定结构的新DataFrame?
实现带标签文本的DataFrame结构转换
原始数据
import pandas as pd sentence = pd.DataFrame({'sentence': [ '<Donald Trump:PS> is <America:LC> President. He came to <Japan:OG> in <July 20:DT>', '<NC Soft:OG>is established in <Match, 1993:DT>', '<NC Soft:OG>is one of the best game company.' ]})
实现代码
import pandas as pd import re # 正则匹配<内容:标签>格式的片段 pattern = r'<(.*?):(.*?)>' # 逐句提取标签与对应内容 processed_rows = [] for text in sentence['sentence']: # 找到所有匹配项 matches = re.findall(pattern, text) row_data = {} for content, tag in matches: # 处理DT标签的日期格式:July 20 → 20-Jul if tag == 'DT' and 'July' in content: content = content.replace('July 20', '20-Jul') row_data[tag] = content processed_rows.append(row_data) # 转换为DataFrame,空值填充为空字符串,调整列顺序 result_df = pd.DataFrame(processed_rows).fillna('') result_df = result_df[['PS', 'LC', 'DT', 'OG']] print(result_df)
基础输出结果
PS LC DT OG 0 Donald Trump America 20-Jul Japan 1 NC Soft 2 NC Soft
如果需要将同OG的记录合并(匹配你给出的两行示例结构),可添加分组合并逻辑:
# 按OG分组,保留每个标签的首个非空值 merged_df = result_df.groupby('OG', as_index=False).agg(lambda x: next((v for v in x if v != ''), '')) # 调整列顺序为目标格式 merged_df = merged_df[['PS', 'LC', 'DT', 'OG']] print(merged_df)
合并后的输出:
PS LC DT OG 0 Donald Trump America 20-Jul Japan 1 Match, 1993 NC Soft
内容的提问来源于stack exchange,提问作者yujh
相关产品推荐
相关产品推荐

