如何基于文本值拆分Pandas DataFrame列并保留含冒号内容?
Pandas文本列拆分:保留含冒号的完整内容
问题场景
现有包含text_col列的Pandas DataFrame:
id text_col 1 Was it Accurate?: Yes\n\nReasoning: This is a sample : text 2 Was it Accurate?: Yes\n\nReasoning: This is a :sample 2 text 3 Was it Accurate?: No\n\nReasoning: This is a sample: 1. text
需要将text_col拆分为Was it Accurate?和Reasoning两列,最终结果需保留Reasoning内容中的所有冒号及换行(如测试用的LLM_response列包含多行内容)。
尝试用"\n\nReasoning:"拆分未达预期,使用正则表达式:
df[['Was it Accurate?', 'Reasoning']] = df['text_col'].str.extract(r'Was it Accurate\?: (Yes|No)\n\nReasoning: (.*)')
时,Reasoning列丢失了冒号后的部分内容,且无法匹配多行文本。
测试用的sample_100数据集字典:
{ 'id_no': [8736215], 'Notes': [' Temp Notes Sample xxxxxxxxxxxxx [4/21/23, 2:10 PM] Work started -work complete-'], 'ProblemDescription': ['Sample problem description xxxxxxxxxxxxxxxxxxxxxxxx'], 'LLM_response': ['Accurate & Understandable: Yes\n\nReasoning: The Technician notes are accurate and understandable as:\n1) The technician provided detailed steps on how they addressed the mold issue by removing materials, treating surfaces, priming, and painting them.\n2) Additionally, even though there was non-repair related information (toilet repairs), the main issue of mold growth was addressed.\n3) The process described logically follows the process for remedying a mold issue, which aligns with the problem description.'], 'Accurate & Understandable': ['Yes'], 'Reasoning': ['The Technician notes are accurate and understandable as:'] }
解决方案
方案1:使用str.split限定拆分次数
利用str.split的n参数只拆分一次,避免内容中的冒号或换行干扰:
# 一步拆分出目标列 df[['temp', 'Reasoning']] = df['text_col'].str.split('\n\nReasoning: ', n=1, expand=True) df[['Was it Accurate?', '_']] = df['temp'].str.split(': ', n=1, expand=True) df = df.drop(['temp', '_'], axis=1)
方案2:修改正则表达式,匹配所有字符(含换行)
添加re.DOTALL标志(或用[\s\S]*代替.*),让正则能匹配包括换行在内的所有内容:
# 方法A:使用re.DOTALL标志 import re df[['Was it Accurate?', 'Reasoning']] = df['text_col'].str.extract(r'Was it Accurate\?: (Yes|No)\n\nReasoning: (.*)', flags=re.DOTALL) # 方法B:用[\s\S]*匹配所有字符(无需额外导入) df[['Was it Accurate?', 'Reasoning']] = df['text_col'].str.extract(r'Was it Accurate\?: (Yes|No)\n\nReasoning: ([\s\S]*)')
针对LLM_response列的类似拆分,只需调整正则前缀即可:
df[['Accurate & Understandable', 'Reasoning']] = df['LLM_response'].str.extract(r'Accurate & Understandable: (Yes|No)\n\nReasoning: ([\s\S]*)')
内容的提问来源于stack exchange,提问作者Shubham R
相关产品推荐
相关产品推荐

