You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将bs4爬取的类字典网页数据转换为标准Python字典

转换方案

有两种高效实现方式,你可以根据自己的场景选择:

方案1:逐行解析(兼容性最优,推荐)

直接拆分原始文本的键值对,同时完成类型转换,不会受特殊字符影响,稳定性最高。
实现代码如下:

def parse_pzwiki_data(raw_str):
    res = {}
    # 去除首尾大括号和全局空白
    content = raw_str.strip().lstrip('{').rstrip('}').strip()
    for line in content.splitlines():
        # 去除行首尾空白、末尾逗号
        clean_line = line.strip().rstrip(',')
        if not clean_line:
            continue
        # 按=拆分键值,最多拆分1次避免值内包含=的异常
        key, val = [part.strip() for part in clean_line.split('=', 1)]
        # 自动转换值的类型
        if val.replace('.', '', 1).isdigit():
            val = float(val) if '.' in val else int(val)
        res[key] = val
    return res

# 调用方式,直接传入你爬取到的data
parsed_result = parse_pzwiki_data(data)

方案2:正则+JSON转换(代码更简洁)

如果你的爬取数据格式非常固定,没有特殊字符,可以用正则快速调整为合法JSON格式后直接转换:

import re
import json

def parse_by_regex_json(raw_str):
    # 给键加双引号,替换=为:,给非数字值加双引号
    processed = re.sub(r'(\w+)\s*=\s*([^\n,]+)', lambda m: f'"{m.group(1)}": "{m.group(2).strip()}"', raw_str)
    # 移除数字值的双引号
    processed = re.sub(r'"(\d+\.?\d*)"', r'\1', processed)
    return json.loads(processed)

# 调用方式
parsed_result = parse_by_regex_json(data)

两种方案处理你给出的示例数据,都能直接得到你期望的标准Python字典格式。

内容的提问来源于stack exchange,提问作者Lun

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.23 15:06:02