如何解析含特殊分隔符的文本文件并转换为DataFrame?
解析特定格式文本为DataFrame的实现方法
输入文本格式
('name: ', u'Jacky') ('male: ', True) ('hobby: ', u'play football and bascket') ('age: ', 24.0) ---------------- ('name: ', u'Belly') ('male: ', True) ('hobby: ', u'dancer') ('age: ', 74.0) ---------------- ('name: ', u'Chow') ('male: ', True) ('hobby: ', u'artist') ('age: ', 46.0)
期望输出DataFrame
name male hobby age 0 Jacky True play football and bascket 24 1 Belly True dancer 74 2 Chow True artist 46
实现代码
下面是用Python和pandas实现的具体方案,直接处理文本格式并转换为DataFrame:
import pandas as pd # 读取目标文本文件 with open('your_file.txt', 'r', encoding='utf-8') as f: raw_content = f.read() # 用分隔符分割每个用户的信息块 user_sections = raw_content.split('----------------') # 初始化列表存储每个用户的信息字典 user_data = [] for section in user_sections: # 跳过空的分割块(避免文件首尾空行影响) if not section.strip(): continue info_dict = {} # 按行处理每个用户的属性 for line in section.strip().split('\n'): # 清理行首尾的括号,分割键值部分 cleaned_line = line.strip().strip('()') key_str, value_str = cleaned_line.split(', ', 1) # 处理键:去掉冒号和多余空格,统一转为小写 clean_key = key_str.strip().rstrip(':').strip().lower() # 处理值:兼容Python2的u前缀,转换布尔/数值类型 clean_value = value_str.strip() # 移除字符串类型的u前缀 if clean_value.startswith(("u'", 'u"')): clean_value = clean_value[2:-1] # 转换布尔值 elif clean_value in ('True', 'False'): clean_value = eval(clean_value) # 转换数值类型(自动区分整数和浮点数) else: try: num_value = float(clean_value) clean_value = int(num_value) if num_value.is_integer() else num_value except ValueError: pass info_dict[clean_key] = clean_value user_data.append(info_dict) # 转换为DataFrame result_df = pd.DataFrame(user_data) print(result_df)
关键处理点说明
- 分割用户块:用
----------------作为分隔符,把整个文本拆分成单个用户的信息组 - 清理键名:去掉键名后的冒号和多余空格,统一转为小写,确保列名规范
- 值类型转换:处理Python2的
u'前缀,自动转换布尔值、整数/浮点数,保证数据类型正确 - 空块过滤:跳过分割后可能出现的空内容块,避免生成无效数据
内容的提问来源于stack exchange,提问作者Michael
相关产品推荐
相关产品推荐

