You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中清洗BeautifulSoup提取的不规则列表生成DataFrame的更优方法

解决方案

核心思路是放弃固定索引取值,改为根据字段标识动态匹配对应值,适配不同长度的子列表,步骤如下:

处理逻辑

  1. 先对每条子列表做基础清洗:过滤空字符串,同时去除所有元素的首尾空格,得到干净的键值交替结构
  2. 日期和id的位置相对固定(清洗后列表的第2、3位,对应索引1、2),直接取值即可
  3. 剩余字段通过遍历匹配字段名(Type/Description/Amount),取对应字段后一位的内容作为值,没有对应字段则默认留空

代码实现

import pandas as pd

souplist = [
['   Date', '', '  Fri 30th Apr 2021', '', ' 60084096-1', 'Type', '', '  Staff Travel (Rail)', '', 'Description', '', '', '  Stratford International', 'Amount', '', '', '', '', '   £25.10 Paid  '],
['   Date', '', '  Tue 27th Apr 2021', '', ' 60084096-3', 'Type', '', '  Office Costs (Stationery & printing)', '', 'Description', '', '', '  AMAZON.CO.UK [***]', 'Amount', '', '', '', '', '   £42.98 Paid  '],
['   Date', '', '  Tue 1st Dec 2020', '', ' 90012371-0', 'Type', '', '  Office Costs (Rent)', '', '  Amount', '', '', '', '', '   £3,500.00 Paid  '],
['   Date', '', '  Wed 14th Oct 2020', '', ' 60064831-1', 'Type', '', '  Office Costs (Software & applications)', '', 'Description', '', '', '  MAILCHIMP', 'MISC', 'Amount', '', '', '', '', '   £38.13 Paid  ']
]

processed_data = []
for row in souplist:
    # 清洗当前行:过滤空值+去首尾空格
    clean_row = [item.strip() for item in row if item.strip()]
    # 初始化行数据,缺失字段默认留空
    row_item = {
        'date': '',
        'id': '',
        'type': '',
        'description': '',
        'amount': ''
    }
    # 取位置固定的date和id
    row_item['date'] = clean_row[1]
    row_item['id'] = clean_row[2]
    # 遍历匹配剩余字段
    for i in range(3, len(clean_row)-1):
        current_key = clean_row[i]
        if current_key == 'Type':
            row_item['type'] = clean_row[i+1]
        elif current_key == 'Description':
            row_item['description'] = clean_row[i+1]
        elif current_key == 'Amount':
            row_item['amount'] = clean_row[i+1]
    processed_data.append(row_item)

# 转换为DataFrame
df = pd.DataFrame(processed_data)

# 可选:对金额字段做格式处理,转为数值类型
df['amount'] = df['amount'].str.replace(r'[£, Paid]', '', regex=True).astype(float)

方案优势

  • 完全适配不同长度的子列表,中间出现额外字段(如示例中的MISC)也不会影响取值
  • 缺失字段自动留空,不会触发索引越界报错
  • 提前做了基础清洗,避免空格、空字符串干扰匹配

内容的提问来源于stack exchange,提问作者user1968419

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 01:06:01