You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将非结构化面包店文本数据转为指定格式DataFrame并保存为CSV

实现方案

你可以通过Python结合Pandas库完成转换,代码如下:

import pandas as pd
import re

# 读取原始TXT文件,请替换为你本地的实际文件路径
with open('bakery_raw.txt', 'r', encoding='utf-8') as f:
    # 预处理:去除每行首尾空白、过滤空行
    lines = [line.strip() for line in f if line.strip()]

result = []
idx = 0
total_lines = len(lines)

while idx < total_lines:
    # 解析面包店名称,统一格式为「Bakery 字母」
    curr_name = lines[idx].lower().replace('bakery', 'Bakery ').strip()
    idx += 1
    # 解析面包店地址
    curr_addr = lines[idx]
    idx += 1
    # 判断是否有可用数据
    next_content = lines[idx]
    if next_content == 'data not available':
        # 无数据的场景,直接添加空记录
        result.append({
            'Name': curr_name,
            'Address': curr_addr,
            'type': pd.NA,
            'salt': pd.NA,
            'sugar': pd.NA,
            'water': pd.NA,
            'flour': pd.NA
        })
        idx += 1
    else:
        # 依次解析3类食品的成分数据
        for _ in range(3):
            type_line = lines[idx]
            # 拆分品类名和数值数组
            product_type, val_str = type_line.split(': ')
            # 提取数组内的四个数值,对应盐、糖、水、面粉
            val_list = list(map(int, re.findall(r'\d+', val_str)))
            salt, sugar, water, flour = val_list
            result.append({
                'Name': curr_name,
                'Address': curr_addr,
                'type': product_type,
                'salt': salt,
                'sugar': sugar,
                'water': water,
                'flour': flour
            })
            idx += 1

# 转换为DataFrame
df = pd.DataFrame(result)
# 导出为CSV文件,encoding使用utf-8-sig兼容Windows Excel打开不乱码
df.to_csv('bakery_final.csv', index=False, encoding='utf-8-sig')

运行上述代码后,生成的DataFrame结构和要求完全一致,导出的CSV文件可直接使用。

内容的提问来源于stack exchange,提问作者karan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 21:27:02