Python中清洗BeautifulSoup提取的不规则列表生成DataFrame的更优方法
解决方案
核心思路是放弃固定索引取值,改为根据字段标识动态匹配对应值,适配不同长度的子列表,步骤如下:
处理逻辑
- 先对每条子列表做基础清洗:过滤空字符串,同时去除所有元素的首尾空格,得到干净的键值交替结构
- 日期和id的位置相对固定(清洗后列表的第2、3位,对应索引1、2),直接取值即可
- 剩余字段通过遍历匹配字段名(Type/Description/Amount),取对应字段后一位的内容作为值,没有对应字段则默认留空
代码实现
import pandas as pd souplist = [ [' Date', '', ' Fri 30th Apr 2021', '', ' 60084096-1', 'Type', '', ' Staff Travel (Rail)', '', 'Description', '', '', ' Stratford International', 'Amount', '', '', '', '', ' £25.10 Paid '], [' Date', '', ' Tue 27th Apr 2021', '', ' 60084096-3', 'Type', '', ' Office Costs (Stationery & printing)', '', 'Description', '', '', ' AMAZON.CO.UK [***]', 'Amount', '', '', '', '', ' £42.98 Paid '], [' Date', '', ' Tue 1st Dec 2020', '', ' 90012371-0', 'Type', '', ' Office Costs (Rent)', '', ' Amount', '', '', '', '', ' £3,500.00 Paid '], [' Date', '', ' Wed 14th Oct 2020', '', ' 60064831-1', 'Type', '', ' Office Costs (Software & applications)', '', 'Description', '', '', ' MAILCHIMP', 'MISC', 'Amount', '', '', '', '', ' £38.13 Paid '] ] processed_data = [] for row in souplist: # 清洗当前行:过滤空值+去首尾空格 clean_row = [item.strip() for item in row if item.strip()] # 初始化行数据,缺失字段默认留空 row_item = { 'date': '', 'id': '', 'type': '', 'description': '', 'amount': '' } # 取位置固定的date和id row_item['date'] = clean_row[1] row_item['id'] = clean_row[2] # 遍历匹配剩余字段 for i in range(3, len(clean_row)-1): current_key = clean_row[i] if current_key == 'Type': row_item['type'] = clean_row[i+1] elif current_key == 'Description': row_item['description'] = clean_row[i+1] elif current_key == 'Amount': row_item['amount'] = clean_row[i+1] processed_data.append(row_item) # 转换为DataFrame df = pd.DataFrame(processed_data) # 可选:对金额字段做格式处理,转为数值类型 df['amount'] = df['amount'].str.replace(r'[£, Paid]', '', regex=True).astype(float)
方案优势
- 完全适配不同长度的子列表,中间出现额外字段(如示例中的MISC)也不会影响取值
- 缺失字段自动留空,不会触发索引越界报错
- 提前做了基础清洗,避免空格、空字符串干扰匹配
内容的提问来源于stack exchange,提问作者user1968419
相关产品推荐
相关产品推荐

