You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python提取指定类别括号内字符串?现有函数存在异常

问题描述

需要从形如Technology (Pyrolysis, Gasification)、Application (Agriculture, Animal Feed, Health & Beauty Products)的文本中,提取指定类别(如Technology、Application)括号()内的内容。

给定示例字符串:

string = ''' U.S. Biochar Market Size, Share & Trends Analysis Report By Technology (Pyrolysis, Gasification), By Application (Agriculture, Animal Feed, Health & Beauty Products), By State, And Segment Forecasts'''

已编写Python函数check_para,但该函数无法适配所有场景,会错误提取无关内容(如By State, And Segment Forecasts),需优化该函数以处理大量同类句式的数据集。

现有函数代码

def check_para(args):
    s = args
    start = s.find('Technology (')
    end = s.find('),', start)
    techno = s[start:end][len('Technology ('):]

    s = args
    start = s.find('Material (')
    end = s.find('),', start)
    mat = s[start:end][len('Material ('):]

    s = args
    start = s.find('Product (')
    end = s.find('),', start)
    prod = s[start:end][len('Product ('):]

    s = args
    start = s.find('Service (')
    end = s.find('),', start)
    serv = s[start:end][len('Service ('):]

    s = args
    start = s.find('Type (')
    end = s.find('),', start)
    typ = s[start:end][len('Type ('):]

    s = args
    start = s.find('Form (')
    end = s.find('),', start)
    form = s[start:end][len('Form ('):]

    s = args
    start = s.find('Application (')
    end = s.find('),', start)
    appli = s[start:end][len('Application ('):]

    s = args
    start = s.find('End Use (')
    end = s.find('),', start)
    enduse = s[start:end][len('End Use ('):]

    s = args
    start = s.find('Derivative Grades (')
    end = s.find('),', start)
    deriv = s[start:end][len('Derivative Grades ('):]

    type1 = deriv + form + typ + serv + prod + techno + mat

    application = appli + enduse

    if len(application) > 0 :
        application = application.replace(', ', '\n')
    else:
        application = 'Application I\nApplication II\nApplication III\n'

    if len(type1) > 0:
        type1 = type1.replace(', ', '\n')
    else:
        type1 = 'Type I\nTypeII\nType III\n'

    return application, type1

优化方案

原函数的核心问题是依赖),作为结束标记,一旦文本格式变化(如括号后无逗号)就会错误截取内容。改用正则表达式可以精准匹配括号内的内容,同时简化代码逻辑:

  • 按用途将类别分组(应用类、类型类)
  • 使用正则模式(类别名称)\s*\((.*?)\)匹配目标内容,非贪婪匹配确保只提取到第一个)为止
  • 批量处理所有目标类别,避免重复的字符串查找操作

优化后的代码

import re

def check_para(text):
    # 定义类别分组
    application_categories = ['Application', 'End Use']
    type_categories = ['Technology', 'Material', 'Product', 'Service', 'Type', 'Form', 'Derivative Grades']
    
    application_items = []
    type_items = []
    
    # 提取应用类内容
    for category in application_categories:
        pattern = re.compile(re.escape(category) + r'\s*\((.*?)\)')
        match = pattern.search(text)
        if match:
            items = match.group(1).split(', ')
            application_items.extend(items)
    
    # 提取类型类内容
    for category in type_categories:
        pattern = re.compile(re.escape(category) + r'\s*\((.*?)\)')
        match = pattern.search(text)
        if match:
            items = match.group(1).split(', ')
            type_items.extend(items)
    
    # 格式化输出,无内容时使用默认值
    application = '\n'.join(application_items) if application_items else 'Application I\nApplication II\nApplication III'
    type1 = '\n'.join(type_items) if type_items else 'Type I\nType II\nType III'
    
    return application, type1

# 测试示例
string = ''' U.S. Biochar Market Size, Share & Trends Analysis Report By Technology (Pyrolysis, Gasification), By Application (Agriculture, Animal Feed, Health & Beauty Products), By State, And Segment Forecasts'''
app, typ = check_para(string)
print("应用类内容:\n", app)
print("\n类型类内容:\n", typ)

优化效果

  • 不再依赖固定的),结束符,适配更多格式场景
  • 代码扩展性强,新增类别只需在对应分组列表中添加
  • 精准提取目标内容,不会错误截取无关文本

内容的提问来源于stack exchange,提问作者Ayush

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 19:15:51