Pandas如何将列表内容转为DataFrame两列并清洗冗余信息
实现方法
处理逻辑如下:
- 先过滤掉所有行首尾去空格后以
-开头的注释行,这类内容都是取值说明类的冗余信息 - 对剩下的有效行,先去除首尾的空格、换行符
- 按第一个出现的冒号
:将字符串拆分为序号部分和指标部分,指标部分再做去空格、移除末尾问号的处理 - 最后将处理好的两列数据转为DataFrame即可
import pandas as pd # 原始列表 sorted_info = [' 1: surgery?\n', ' 2: Age\n', ' 3: Hospital Number\n', ' 4: rectal temperature\n', ' 5: pulse\n', ' - is a reflection of the heart condition: 30 -40 is normal for adults\n', ' 6: respiratory rate\n', ' 7: temperature of extremities\n', ' - possible values:\n', ' 8: peripheral pulse\n', ' - possible values are:\n', ' 9: mucous membranes\n', ' - possible values are:\n', ' 10: capillary refill time\n', " 11: pain - a subjective judgement of the horse's pain level\n", ' - possible values:\n', ' 12: peristalsis\n', ' - possible values:\n', ' 13: abdominal distension\n', ' 14: nasogastric tube\n', ' - possible values:\n', ' 15: nasogastric reflux\n', ' 16: nasogastric reflux PH\n', ' 17: rectal examination - feces\n', ' 18: abdomen\n', ' 19: packed cell volume\n', ' 20: total protein\n', ' 21: abdominocentesis appearance\n', ' - possible values:\n', ' 22: abdomcentesis total protein\n', ' 23: outcome\n', ' - possible values:\n', ' 24: surgical lesion?\n', ' - possible values:\n', ' 25, 26, 27: type of lesion\n', ' 28: cp_data\n'] processed_data = [] for line in sorted_info: clean_line = line.strip() # 过滤冗余注释行 if clean_line.startswith('-'): continue # 拆分序号和指标 colon_pos = clean_line.find(':') col1 = clean_line[:colon_pos + 1] col2 = clean_line[colon_pos + 1:].strip().rstrip('?') processed_data.append([col1, col2]) # 生成符合要求的DataFrame df = pd.DataFrame(processed_data, columns=['Col1', 'Col2']) print(df)
运行代码后输出的DataFrame完全匹配你期望的结构,所有冗余信息都会被清理,指标内容会按要求拆分。
内容的提问来源于stack exchange,提问作者This_guy
相关产品推荐
相关产品推荐

