如何用Python提取指定类别括号内字符串?现有函数存在异常
问题描述
需要从形如Technology (Pyrolysis, Gasification)、Application (Agriculture, Animal Feed, Health & Beauty Products)的文本中,提取指定类别(如Technology、Application)括号()内的内容。
给定示例字符串:
string = ''' U.S. Biochar Market Size, Share & Trends Analysis Report By Technology (Pyrolysis, Gasification), By Application (Agriculture, Animal Feed, Health & Beauty Products), By State, And Segment Forecasts'''
已编写Python函数check_para,但该函数无法适配所有场景,会错误提取无关内容(如By State, And Segment Forecasts),需优化该函数以处理大量同类句式的数据集。
现有函数代码
def check_para(args): s = args start = s.find('Technology (') end = s.find('),', start) techno = s[start:end][len('Technology ('):] s = args start = s.find('Material (') end = s.find('),', start) mat = s[start:end][len('Material ('):] s = args start = s.find('Product (') end = s.find('),', start) prod = s[start:end][len('Product ('):] s = args start = s.find('Service (') end = s.find('),', start) serv = s[start:end][len('Service ('):] s = args start = s.find('Type (') end = s.find('),', start) typ = s[start:end][len('Type ('):] s = args start = s.find('Form (') end = s.find('),', start) form = s[start:end][len('Form ('):] s = args start = s.find('Application (') end = s.find('),', start) appli = s[start:end][len('Application ('):] s = args start = s.find('End Use (') end = s.find('),', start) enduse = s[start:end][len('End Use ('):] s = args start = s.find('Derivative Grades (') end = s.find('),', start) deriv = s[start:end][len('Derivative Grades ('):] type1 = deriv + form + typ + serv + prod + techno + mat application = appli + enduse if len(application) > 0 : application = application.replace(', ', '\n') else: application = 'Application I\nApplication II\nApplication III\n' if len(type1) > 0: type1 = type1.replace(', ', '\n') else: type1 = 'Type I\nTypeII\nType III\n' return application, type1
优化方案
原函数的核心问题是依赖),作为结束标记,一旦文本格式变化(如括号后无逗号)就会错误截取内容。改用正则表达式可以精准匹配括号内的内容,同时简化代码逻辑:
- 按用途将类别分组(应用类、类型类)
- 使用正则模式
(类别名称)\s*\((.*?)\)匹配目标内容,非贪婪匹配确保只提取到第一个)为止 - 批量处理所有目标类别,避免重复的字符串查找操作
优化后的代码
import re def check_para(text): # 定义类别分组 application_categories = ['Application', 'End Use'] type_categories = ['Technology', 'Material', 'Product', 'Service', 'Type', 'Form', 'Derivative Grades'] application_items = [] type_items = [] # 提取应用类内容 for category in application_categories: pattern = re.compile(re.escape(category) + r'\s*\((.*?)\)') match = pattern.search(text) if match: items = match.group(1).split(', ') application_items.extend(items) # 提取类型类内容 for category in type_categories: pattern = re.compile(re.escape(category) + r'\s*\((.*?)\)') match = pattern.search(text) if match: items = match.group(1).split(', ') type_items.extend(items) # 格式化输出,无内容时使用默认值 application = '\n'.join(application_items) if application_items else 'Application I\nApplication II\nApplication III' type1 = '\n'.join(type_items) if type_items else 'Type I\nType II\nType III' return application, type1 # 测试示例 string = ''' U.S. Biochar Market Size, Share & Trends Analysis Report By Technology (Pyrolysis, Gasification), By Application (Agriculture, Animal Feed, Health & Beauty Products), By State, And Segment Forecasts''' app, typ = check_para(string) print("应用类内容:\n", app) print("\n类型类内容:\n", typ)
优化效果
- 不再依赖固定的
),结束符,适配更多格式场景 - 代码扩展性强,新增类别只需在对应分组列表中添加
- 精准提取目标内容,不会错误截取无关文本
内容的提问来源于stack exchange,提问作者Ayush
相关产品推荐
相关产品推荐

