You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解析含特殊分隔符的文本文件并转换为DataFrame?

解析特定格式文本为DataFrame的实现方法

输入文本格式

('name:   ', u'Jacky')
('male:   ', True)
('hobby:   ', u'play football and bascket')
('age:   ', 24.0)
----------------
('name:   ', u'Belly')
('male:   ', True)
('hobby:   ', u'dancer')
('age:   ', 74.0)
----------------
('name:   ', u'Chow')
('male:   ', True)
('hobby:   ', u'artist')
('age:   ', 46.0)

期望输出DataFrame

name  male                     hobby  age
0  Jacky  True  play football and bascket   24
1  Belly  True                     dancer   74
2   Chow  True                     artist   46

实现代码

下面是用Python和pandas实现的具体方案,直接处理文本格式并转换为DataFrame:

import pandas as pd

# 读取目标文本文件
with open('your_file.txt', 'r', encoding='utf-8') as f:
    raw_content = f.read()

# 用分隔符分割每个用户的信息块
user_sections = raw_content.split('----------------')

# 初始化列表存储每个用户的信息字典
user_data = []
for section in user_sections:
    # 跳过空的分割块(避免文件首尾空行影响)
    if not section.strip():
        continue
    info_dict = {}
    # 按行处理每个用户的属性
    for line in section.strip().split('\n'):
        # 清理行首尾的括号,分割键值部分
        cleaned_line = line.strip().strip('()')
        key_str, value_str = cleaned_line.split(', ', 1)
        
        # 处理键:去掉冒号和多余空格,统一转为小写
        clean_key = key_str.strip().rstrip(':').strip().lower()
        
        # 处理值:兼容Python2的u前缀,转换布尔/数值类型
        clean_value = value_str.strip()
        # 移除字符串类型的u前缀
        if clean_value.startswith(("u'", 'u"')):
            clean_value = clean_value[2:-1]
        # 转换布尔值
        elif clean_value in ('True', 'False'):
            clean_value = eval(clean_value)
        # 转换数值类型(自动区分整数和浮点数)
        else:
            try:
                num_value = float(clean_value)
                clean_value = int(num_value) if num_value.is_integer() else num_value
            except ValueError:
                pass
        
        info_dict[clean_key] = clean_value
    user_data.append(info_dict)

# 转换为DataFrame
result_df = pd.DataFrame(user_data)
print(result_df)

关键处理点说明

  • 分割用户块:用----------------作为分隔符,把整个文本拆分成单个用户的信息组
  • 清理键名:去掉键名后的冒号和多余空格,统一转为小写,确保列名规范
  • 值类型转换:处理Python2的u'前缀,自动转换布尔值、整数/浮点数,保证数据类型正确
  • 空块过滤:跳过分割后可能出现的空内容块,避免生成无效数据

内容的提问来源于stack exchange,提问作者Michael

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 13:57:28