Python动态添加字典键后导出CSV报错,如何修正数据结构?
解决Python爬虫动态字典键的CSV导出与JSON结构问题
问题根源
- JSON结构错误:原代码遍历价格详情时,每次循环都创建新字典并添加到结果列表,导致每个
lowest_price_{idx}对应独立条目,而非合并到同一个产品字典中。 - CSV导出报错:
csv.DictWriter使用第一个字典的键作为表头,后续字典包含lowest_price_1等新键时,因不在初始表头集合中抛出ValueError。
修复步骤
1. 合并动态键到单个产品字典
为每个产品先创建基础字典,再遍历价格列表将动态键逐个添加到该字典,最后将完整字典加入结果列表,避免重复创建条目。
2. 动态生成CSV表头
收集所有结果字典的键并去重,确保覆盖所有动态生成的键,避免表头缺失。
修改后的完整代码
import requests from bs4 import BeautifulSoup import csv class ZiwiScraper: results = [] headers = { 'authority': '99petshops.com.au', 'accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.9', 'accept-language': 'en,ru;q=0.9', 'cache-control': 'max-age=0', 'referer': 'https://www.upwork.com/', 'sec-ch-ua': '"Chromium";v="104", " Not A;Brand";v="99", "Yandex";v="22"', 'sec-ch-ua-mobile': '?0', 'sec-ch-ua-platform': '"Linux"', 'sec-fetch-dest': 'document', 'sec-fetch-mode': 'navigate', 'sec-fetch-site': 'cross-site', 'sec-fetch-user': '?1', 'upgrade-insecure-requests': '1', 'user-agent': 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/104.0.5112.114 YaBrowser/22.9.1.1110 (beta) Yowser/2.5 Safari/537.36', } def fetch(self, url): print(f'HTTP GET request to URL: {url}', end='') res = requests.get(url, headers=self.headers) print(f' | Status Code: {res.status_code}') return res def parse(self, html): soup = BeautifulSoup(html, 'lxml') titles = [title.text.strip() for title in soup.find_all('h2')] low_prices = [low_price.text.split(' ')[-1] for low_price in soup.find_all('span', {'class': 'hilighted'})] store_names = [] stores = soup.find_all('p') for store in stores: store_name = store.find('img') if store_name: store_names.append(store_name['alt']) shipping_prices = [shipping.text.strip() for shipping in soup.find_all('p', {'class': 'shipping'})] price_per_hundered_kg = [unit_per_kg.text.strip() for unit_per_kg in soup.find_all('p', {'class': 'unit-price'})] other_details = soup.find_all('div', {'class': 'pd-details'}) for index in range(len(titles)): # 初始化单个产品的基础字典 product_dict = { 'title': titles[index], 'lowest_prices': low_prices[index] if index < len(low_prices) else '', 'store_names': store_names[index] if index < len(store_names) else '', 'shipping_prices': shipping_prices[index] if index < len(shipping_prices) else '', 'price_per_100_kg': price_per_hundered_kg[index] if index < len(price_per_hundered_kg) else '', } # 为当前产品添加所有动态价格键 if index < len(other_details): detail = other_details[index] detail_1 = [pr.text.strip() for pr in detail.find_all('span', {'class': 'sp-price'})] for idx, price in enumerate(detail_1): product_dict[f'lowest_price_{idx}'] = price # 将完整产品字典加入结果列表 self.results.append(product_dict) def to_csv(self): if not self.results: print('No results to export') return # 收集所有可能的字段名,覆盖动态生成的键 fieldnames = set() for row in self.results: fieldnames.update(row.keys()) fieldnames = sorted(fieldnames) # 排序保证表头顺序一致 with open('ziwi_pets_2.csv', 'w', newline='') as csv_file: writer = csv.DictWriter(csv_file, fieldnames=fieldnames) writer.writeheader() for row in self.results: writer.writerow(row) print('Stored results to "ziwi_pets_2.csv"') def run(self): for page in range(1): url = f'https://99petshops.com.au/Search?brandName=Ziwi%20Peak&animalCode=DOG&storeId=89%2F&page={page}' response = self.fetch(url) self.parse(response.text) self.to_csv() if __name__ == '__main__': scraper = ZiwiScraper() scraper.run()
修复后效果
- JSON结构:每个产品对应一个字典,包含所有
lowest_price_0、lowest_price_1等动态键,与预期格式一致。 - CSV导出:不再报错,所有动态键都会被包含在表头中,完整导出所有爬取数据。
内容的提问来源于stack exchange,提问作者X-something
相关产品推荐
相关产品推荐

