You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python动态添加字典键后导出CSV报错,如何修正数据结构?

解决Python爬虫动态字典键的CSV导出与JSON结构问题

问题根源

  1. JSON结构错误:原代码遍历价格详情时,每次循环都创建新字典并添加到结果列表,导致每个lowest_price_{idx}对应独立条目,而非合并到同一个产品字典中。
  2. CSV导出报错:csv.DictWriter使用第一个字典的键作为表头,后续字典包含lowest_price_1等新键时,因不在初始表头集合中抛出ValueError。

修复步骤

1. 合并动态键到单个产品字典

为每个产品先创建基础字典,再遍历价格列表将动态键逐个添加到该字典,最后将完整字典加入结果列表,避免重复创建条目。

2. 动态生成CSV表头

收集所有结果字典的键并去重,确保覆盖所有动态生成的键,避免表头缺失。

修改后的完整代码

import requests
from bs4 import BeautifulSoup
import csv

class ZiwiScraper:
    results = []
    
    headers = {
        'authority': '99petshops.com.au',
        'accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.9',
        'accept-language': 'en,ru;q=0.9',
        'cache-control': 'max-age=0',
        'referer': 'https://www.upwork.com/',
        'sec-ch-ua': '"Chromium";v="104", " Not A;Brand";v="99", "Yandex";v="22"',
        'sec-ch-ua-mobile': '?0',
        'sec-ch-ua-platform': '"Linux"',
        'sec-fetch-dest': 'document',
        'sec-fetch-mode': 'navigate',
        'sec-fetch-site': 'cross-site',
        'sec-fetch-user': '?1',
        'upgrade-insecure-requests': '1',
        'user-agent': 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/104.0.5112.114 YaBrowser/22.9.1.1110 (beta) Yowser/2.5 Safari/537.36',
    }

    def fetch(self, url):
        print(f'HTTP GET request to URL: {url}', end='')
        res = requests.get(url, headers=self.headers)
        print(f' | Status Code: {res.status_code}')
        return res
    
    def parse(self, html):
        soup = BeautifulSoup(html, 'lxml')
        titles = [title.text.strip() for title in soup.find_all('h2')]
        low_prices = [low_price.text.split(' ')[-1] for low_price in soup.find_all('span', {'class': 'hilighted'})]
        store_names = []
        stores = soup.find_all('p')
        for store in stores:
            store_name = store.find('img')
            if store_name:
                store_names.append(store_name['alt'])
        shipping_prices = [shipping.text.strip() for shipping in soup.find_all('p', {'class': 'shipping'})]
        price_per_hundered_kg = [unit_per_kg.text.strip() for unit_per_kg in soup.find_all('p', {'class': 'unit-price'})]
        other_details = soup.find_all('div', {'class': 'pd-details'})
        
        for index in range(len(titles)):
            # 初始化单个产品的基础字典
            product_dict = {
                'title': titles[index],
                'lowest_prices': low_prices[index] if index < len(low_prices) else '',
                'store_names': store_names[index] if index < len(store_names) else '',
                'shipping_prices': shipping_prices[index] if index < len(shipping_prices) else '',
                'price_per_100_kg': price_per_hundered_kg[index] if index < len(price_per_hundered_kg) else '',
            }
            
            # 为当前产品添加所有动态价格键
            if index < len(other_details):
                detail = other_details[index]
                detail_1 = [pr.text.strip() for pr in detail.find_all('span', {'class': 'sp-price'})]
                for idx, price in enumerate(detail_1):
                    product_dict[f'lowest_price_{idx}'] = price
            
            # 将完整产品字典加入结果列表
            self.results.append(product_dict)
            
    def to_csv(self):
        if not self.results:
            print('No results to export')
            return
        
        # 收集所有可能的字段名,覆盖动态生成的键
        fieldnames = set()
        for row in self.results:
            fieldnames.update(row.keys())
        fieldnames = sorted(fieldnames)  # 排序保证表头顺序一致
        
        with open('ziwi_pets_2.csv', 'w', newline='') as csv_file:
            writer = csv.DictWriter(csv_file, fieldnames=fieldnames)
            writer.writeheader()
            for row in self.results:
                writer.writerow(row)
            print('Stored results to "ziwi_pets_2.csv"')
            
    def run(self):
        for page in range(1):
            url = f'https://99petshops.com.au/Search?brandName=Ziwi%20Peak&animalCode=DOG&storeId=89%2F&page={page}'
            response = self.fetch(url)
            self.parse(response.text)
        self.to_csv()

if __name__ == '__main__':
    scraper = ZiwiScraper()
    scraper.run()

修复后效果

  1. JSON结构:每个产品对应一个字典,包含所有lowest_price_0、lowest_price_1等动态键,与预期格式一致。
  2. CSV导出:不再报错,所有动态键都会被包含在表头中,完整导出所有爬取数据。

内容的提问来源于stack exchange,提问作者X-something

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 13:50:28