如何用正则从Pandas DataFrame抽取指定内容生成新列
正则抽取修正方案
错误原因
你原有的正则r'Thrown: lib: (.*(?:\r?\n.*)*)'会捕获Thrown: lib:之后的全部内容,不符合「仅匹配到第一个换行前」的需求。
完整实现代码
import pandas as pd # 初始化DataFrame data = {'c1':['Level: LOGGING_ONLY\n Thrown: lib: this is problem type 01\n \t\n \tError executing the statement: error statement 1\n', 'Level: NOT_LOGGING_ONLY\n Thrown: lib: this is problem type 01\n \t\n \tError executing the statement: error statement 3\n', 'Level: LOGGING_ONLY\n Thrown: lib: this is problem type 02\n \t\n \tError executing the statement: error statement2\n', 'Level: NOT_LOGGING_ONLY\n Thrown: lib: this is problem type 04\n \t\n \tError executing the statement: error statement1\n'], 'c2':["one", "two", "three", "four"]} df = pd.DataFrame(data) # 抽取c3(group 01):匹配Thrown: lib:后到第一个换行前的内容,自动去除首尾多余空格 df['c3'] = df['c1'].str.extract(r'Thrown: lib:\s*(.*?)\s*\n', expand=False) # 抽取c4:跳过group 01后所有换行、制表符、空白字符,匹配到下一个换行前的内容 df['c4'] = df['c1'].str.extract(r'Thrown: lib:.*?\n[\s\t\n]*(.*?)\s*\n', expand=False) # 调整列顺序为你需要的最终结构 df = df[['c1', 'c3', 'c4', 'c2']]
正则说明
- 抽取c3的正则
r'Thrown: lib:\s*(.*?)\s*\n':\s*匹配Thrown: lib:后可能存在的多余空格.*?为非贪婪匹配,会在遇到第一个\n时停止匹配- 尾部
\s*会自动去除捕获内容末尾的多余空格,不需要额外做清洗
- 抽取c4的正则
r'Thrown: lib:.*?\n[\s\t\n]*(.*?)\s*\n':- 先匹配到group 01的末尾换行
[\s\t\n]*跳过两段内容之间所有的空白、换行、制表符- 同样用非贪婪匹配拿到下一个换行前的有效内容
内容的提问来源于stack exchange,提问作者Catarina Nogueira
相关产品推荐
相关产品推荐

