You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何访问嵌套字典列表并修复手机存储、状态提取报错及导出CSV

手机数据提取错误修复与CSV导出方案

问题背景

从data列表提取手机存储(如4GB)和状态(全新/二手)时触发错误:

AttributeError: 'list' object has no attribute 'values'

需修复错误并将最终数据导出为CSV文件。以下是data列表结构及原代码:

data列表结构示例

[
    {
        'admin_info': {}, 'as_top': False,
        'attrs': [
            {'name': 'Condition', 'value': 'Used', 'unit': None},
            {'name': 'Screen Size', 'value': '4-5 inches', 'unit': None},
            {'name':  'RAM', 'value': '3 GB', 'unit': None}
        ],
        'price_obj': {'value': 215000, 'view': '₦ 215,000', 'period': None},
        'region_name': 'Akure',
        'region_parent_name': 'Ondo State',
        'title': 'Apple iPhone XS Max 64 GB Gold',
        # 其他字段省略...
    }
]

原代码

import requests
from bs4 import BeautifulSoup

# list to hold all results
data = []

# for multiple pages iterate in range from - to
for i in range(0, 40):
    # increase page by number of iteration
    url = f'https://jiji.ng/api_web/v1/listing?slug=mobile-phones&init_page=true&page={i}'
    # extend list with all items per page
    data.extend(requests.get(url).json()['adverts_list']['adverts'])

    for name in data:
        name = name.get('title')
    for location in data:
        location = location.get('region_name')
    for state in data:
        state = state.get('region_parent_name')
    for price in data:
        price = price.get('price_obj')['value']
    for storage in data:
        storage = storage.get('attrs').values()
    for status in data:
        status = status.get('attrs')[0][2]
        print([name, location, state, price, storage, status])

错误分析

  1. attrs.values()错误:attrs是列表类型,每个元素是字典,不能直接调用字典的.values()方法。
  2. 状态取值错误:attrs中的元素是字典,需通过键名'value'取值,而非索引[0][2]。
  3. 循环逻辑错误:多个独立for循环遍历整个data列表,最终变量仅保留最后一个元素的值,无法逐个处理广告项。

修复后代码(含CSV导出)

import requests
import csv

# 存储所有爬取到的广告数据
data = []

# 爬取40页手机广告数据
for i in range(0, 40):
    url = f'https://jiji.ng/api_web/v1/listing?slug=mobile-phones&init_page=true&page={i}'
    try:
        response = requests.get(url).json()
        data.extend(response['adverts_list']['adverts'])
    except Exception as e:
        print(f"第{i}页爬取失败: {e}")
        continue

# 整理提取结构化数据
processed_data = []
for item in data:
    # 提取基础字段,添加默认值避免KeyError
    title = item.get('title', 'N/A')
    location = item.get('region_name', 'N/A')
    state = item.get('region_parent_name', 'N/A')
    price = item.get('price_obj', {}).get('value', 'N/A')
    
    # 从attrs中提取状态和存储信息
    condition = 'N/A'
    storage = 'N/A'
    for attr in item.get('attrs', []):
        attr_name = attr.get('name')
        if attr_name == 'Condition':
            condition = attr.get('value')
        # 适配RAM/Storage两种可能的字段名
        elif attr_name in ['RAM', 'Storage']:
            storage = attr.get('value')
    
    processed_data.append([title, location, state, price, storage, condition])

# 导出到CSV文件
with open('mobile_phones.csv', 'w', newline='', encoding='utf-8') as csv_file:
    writer = csv.writer(csv_file)
    # 写入表头
    writer.writerow(['标题', '地区', '州', '价格', '存储/内存', '状态'])
    # 写入所有数据
    writer.writerows(processed_data)

print("数据已成功导出至mobile_phones.csv")

代码改进说明

  • 单个循环遍历每个广告项,确保每条数据都被正确提取。
  • 遍历attrs列表匹配字段名,避免依赖固定索引(字段顺序可能变化)。
  • 添加异常处理和默认值,避免因字段缺失导致程序崩溃。
  • 使用Python内置csv模块实现标准CSV导出,兼容各类表格工具。

内容的提问来源于stack exchange,提问作者Kelly

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 16:37:18