You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Beautiful Soup提取网页JSON中display_name值遇KeyError的解决方法

解决Beautiful Soup提取JSON中display_name并保存CSV的问题

看起来你现在卡在遍历JSON结构提取display_name的环节,还遇到了KeyError,我来帮你梳理下问题出在哪,以及怎么解决:

问题根源

你当前的代码直接用jsons[0]['display_name'],但看你给出的JSON结构,外层是一个包含page_size和refinements的字典,根本没有索引为0的元素,这就是KeyError的原因。而且你需要提取的是两层display_name:一是refinements数组里每个对象的display_name,二是每个refinements下values数组里的display_name。

修正步骤及代码

1. 正确解析JSON内容

首先要把Beautiful Soup拿到的script标签文本,转换成Python可遍历的字典。注意你的示例JSON用的是单引号,而json.loads只识别双引号,所以需要先替换:

import json
from bs4 import BeautifulSoup
import csv

# 假设你已经拿到了soup2对象
for productoslvl in soup2.findAll('script', {'id': 'searchResult'}):
    # 提取script标签里的文本,替换单引号为双引号(处理格式问题)
    json_text = productoslvl.string.replace("'", "\"")
    # 解析为Python字典
    data = json.loads(json_text)

2. 遍历所有display_name

现在我们可以遍历refinements数组,分别提取两层的display_name:

# 用来收集所有display_name的列表
    all_display_names = []

    # 遍历refinements里的每个分类
    for refinement in data['refinements']:
        # 添加分类本身的display_name
        all_display_names.append(refinement['display_name'])
        # 遍历该分类下values里的每个选项
        for value in refinement['values']:
            all_display_names.append(value['display_name'])

3. 保存到CSV文件

最后把收集到的内容写入CSV:

# 写入CSV
    with open('display_names.csv', 'w', newline='', encoding='utf-8') as csvfile:
        writer = csv.writer(csvfile)
        # 写入表头(可选)
        writer.writerow(['display_name'])
        # 逐个写入每个display_name
        for name in all_display_names:
            writer.writerow([name])

完整代码整合

import json
from bs4 import BeautifulSoup
import csv

# 假设这里是你获取soup2的代码,比如:
# with open('your_page.html', 'r', encoding='utf-8') as f:
#     soup2 = BeautifulSoup(f, 'html.parser')

for productoslvl in soup2.findAll('script', {'id': 'searchResult'}):
    json_text = productoslvl.string.replace("'", "\"")
    try:
        data = json.loads(json_text)
    except json.JSONDecodeError as e:
        print(f"JSON解析出错: {e}")
        continue
    
    all_display_names = []
    for refinement in data['refinements']:
        all_display_names.append(refinement['display_name'])
        for value in refinement['values']:
            all_display_names.append(value['display_name'])
    
    with open('display_names.csv', 'w', newline='', encoding='utf-8') as csvfile:
        writer = csv.writer(csvfile)
        writer.writerow(['display_name'])
        for name in all_display_names:
            writer.writerow([name])

额外提示

  • 如果原始JSON的格式不是单引号,就不用做替换操作,直接解析即可
  • 加了try-except捕获JSON解析错误,避免因为格式问题导致程序崩溃
  • 保存CSV时指定encoding='utf-8',防止中文等特殊字符乱码

内容的提问来源于stack exchange,提问作者Rocko

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 06:44:52