You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python提取关键词后的指定文本并整理为字典结构

问题:提取车辆信息并整理为指定格式的Python字典

原始文本内容

# car name 
BMW
suzuki

# car model 
X1 
TT

# color
red 
blue

当前使用的Python代码

keywords = [car_name,car_model,color]
parsed_content = {}

def car_info(text):
    content = {}
    indices = []
    keys = []
    for key in Keywords:
        try:
            content[key] = text[text.index(key) + len(key):]
            indices.append(text.index(key))
            keys.append(key)
        except:
            pass         
    zipped_lists = zip(indices, keys)
    sorted_pairs = sorted(zipped_lists)
    # sorted_pairs

    tuples = zip(*sorted_pairs)
    indices, keys = [ list(tuple) for tuple in  tuples]
    # return keys
    print(keys)

    content = []
    for idx in range(len(indices)):
        if idx != len(indices)-1:
            content.append(text[indices[idx]: indices[idx+1]])
        else:
            content.append(text[indices[idx]: ])
        
    for i in range(len(indices)):
        parsed_content[keys[i]] = content[i]   
    return parsed_content

当前代码输出

parsed_content = {car_name : car_name BMW SUZUKI,
car_model : car_model x1 tt,
color : color red blue

 }

期望输出

{'car_name': ['bmw', 'suzuki'],
 'car_model': ['x1', 'TT'],
 'color': ['red', 'blue']
}

修改方案

原代码未正确分割关键词对应内容,也未去除无效信息、转换格式。以下是修正后的代码:

def car_info(text):
    # 拆分文本为行,过滤空行并去除每行前后空白
    lines = [line.strip() for line in text.splitlines() if line.strip()]
    parsed_content = {}
    current_key = None
    
    for line in lines:
        # 识别标题行(以#开头)
        if line.startswith('#'):
            # 提取并格式化字典键:去掉#、空格,替换空格为下划线
            current_key = line.replace('#', '').strip().replace(' ', '_')
            parsed_content[current_key] = []
        else:
            # 将内容加入当前键对应的列表,car_name下内容统一转小写
            if current_key is not None:
                parsed_content[current_key].append(line.lower() if current_key == 'car_name' else line)
    
    return parsed_content

# 测试用文本
text = """# car name 
BMW
suzuki

# car model 
X1 
TT

# color
red 
blue"""

# 调用并打印结果
result = car_info(text)
print(result)

代码说明

  1. 预处理文本:拆分后过滤空行,避免无效内容干扰。
  2. 自动识别关键词:通过标题行自动生成字典键,无需硬编码关键词列表,扩展性更强。
  3. 整理内容格式:将每个标题下的内容转为列表,car_name字段统一转小写(不需要可移除lower())。

运行后输出与期望一致:

{'car_name': ['bmw', 'suzuki'], 'car_model': ['X1', 'TT'], 'color': ['red', 'blue']}

内容的提问来源于stack exchange,提问作者endpoints

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 05:50:37