如何用Python提取关键词后的指定文本并整理为字典结构
问题:提取车辆信息并整理为指定格式的Python字典
原始文本内容
# car name BMW suzuki # car model X1 TT # color red blue
当前使用的Python代码
keywords = [car_name,car_model,color] parsed_content = {} def car_info(text): content = {} indices = [] keys = [] for key in Keywords: try: content[key] = text[text.index(key) + len(key):] indices.append(text.index(key)) keys.append(key) except: pass zipped_lists = zip(indices, keys) sorted_pairs = sorted(zipped_lists) # sorted_pairs tuples = zip(*sorted_pairs) indices, keys = [ list(tuple) for tuple in tuples] # return keys print(keys) content = [] for idx in range(len(indices)): if idx != len(indices)-1: content.append(text[indices[idx]: indices[idx+1]]) else: content.append(text[indices[idx]: ]) for i in range(len(indices)): parsed_content[keys[i]] = content[i] return parsed_content
当前代码输出
parsed_content = {car_name : car_name BMW SUZUKI, car_model : car_model x1 tt, color : color red blue }
期望输出
{'car_name': ['bmw', 'suzuki'], 'car_model': ['x1', 'TT'], 'color': ['red', 'blue'] }
修改方案
原代码未正确分割关键词对应内容,也未去除无效信息、转换格式。以下是修正后的代码:
def car_info(text): # 拆分文本为行,过滤空行并去除每行前后空白 lines = [line.strip() for line in text.splitlines() if line.strip()] parsed_content = {} current_key = None for line in lines: # 识别标题行(以#开头) if line.startswith('#'): # 提取并格式化字典键:去掉#、空格,替换空格为下划线 current_key = line.replace('#', '').strip().replace(' ', '_') parsed_content[current_key] = [] else: # 将内容加入当前键对应的列表,car_name下内容统一转小写 if current_key is not None: parsed_content[current_key].append(line.lower() if current_key == 'car_name' else line) return parsed_content # 测试用文本 text = """# car name BMW suzuki # car model X1 TT # color red blue""" # 调用并打印结果 result = car_info(text) print(result)
代码说明
- 预处理文本:拆分后过滤空行,避免无效内容干扰。
- 自动识别关键词:通过标题行自动生成字典键,无需硬编码关键词列表,扩展性更强。
- 整理内容格式:将每个标题下的内容转为列表,
car_name字段统一转小写(不需要可移除lower())。
运行后输出与期望一致:
{'car_name': ['bmw', 'suzuki'], 'car_model': ['X1', 'TT'], 'color': ['red', 'blue']}
内容的提问来源于stack exchange,提问作者endpoints
相关产品推荐
相关产品推荐

