非分页站点Selenium+BS4爬取JSON为空的修复方案咨询
问题修复:Selenium+BS4爬取数据生成空JSON的解决方法
问题诊断
当前代码生成空JSON对象的核心原因:
- 循环内重复调用
driver.get(url),导致每次循环都重新加载第一页,无法实现有效翻页,且干扰页面数据渲染 - 仅等待表格元素存在,未等待表格内的具体数据行加载完成,BeautifulSoup解析时无法定位到
<tr align='left'>等有效元素 - 缺失旧代码中价格、零售信息等关键字段的提取逻辑
修复方案(不破坏原有代码结构)
1. 修正翻页逻辑,移除循环内的重复页面加载
将driver.get(url)移到循环外部,仅在初始时加载一次页面,后续通过点击「Next」按钮实现翻页,避免重复回到第一页。
2. 优化等待条件,确保数据完全渲染
将等待表格存在改为等待表格内的数据行可见,确保页面内容完全加载后再解析:
WebDriverWait(driver, 10).until(EC.visibility_of_element_located((By.CSS_SELECTOR, "#ContentPlaceHolder1_ctlSets_GridViewSets tr[align='left']")))
3. 补全字段提取逻辑
复用旧代码中关于价格、零售信息等字段的提取逻辑,确保数据完整。
4. 翻页后增加等待,确保新页面加载完成
点击「Next」按钮后,等待新页面的数据行加载完成,再进入下一次循环解析。
完整修复代码
import json import time from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup # Initialize WebDriver (Safari, Chrome, Firefox, etc.) driver = webdriver.Chrome() # or change to webdriver.Firefox() or webdriver.Safari() url = "https://www.brickeconomy.com/sets/year/2024" max_iterations = 2 # Specify how many pages to fetch delay_seconds = 2 # Delay time between each page transition (seconds) all_sets_data = [] # List to hold all set data try: # 仅初始加载一次页面 driver.get(url) for i in range(max_iterations): # 等待表格内的数据行可见,确保内容完全渲染 WebDriverWait(driver, 10).until(EC.visibility_of_element_located((By.CSS_SELECTOR, "#ContentPlaceHolder1_ctlSets_GridViewSets tr[align='left']"))) # Process the HTML content of the page using BeautifulSoup soup = BeautifulSoup(driver.page_source, 'html.parser') sets_data = [] # Find all rows in the table table = soup.find('table', id='ContentPlaceHolder1_ctlSets_GridViewSets') if table: table_rows = table.find_all('tr', align='left') # Extract set information from each row for row in table_rows: set_info = {} # Find the <h4> element containing the set name and ID set_name_elem = row.find('h4') if set_name_elem: set_string = set_name_elem.text.strip() set_info['id'], set_info['name'] = set_string.split(' ', 1) # Find <div> elements containing Year, Pieces/Minifigs, and other information div_elements = row.find_all('div', class_='mb-2') for div in div_elements: label = div.find('small', class_='text-muted mr-5') if label: label_text = label.text.strip() if label_text == 'Year': set_info['year'] = div.text.replace('Year', '').strip() # 补全价格、零售信息提取逻辑(复用旧代码) td_elements = row.find_all('td', class_='ctlsets-right text-right') for td in td_elements: div_elements = td.find_all('div') for div in div_elements: if "Retail" in div.text: retail_price = div.text.strip() price_without_retail = ' '.join(retail_price.split()[1:]) set_info['price'] = price_without_retail first_sibling = div.find_next_sibling() if first_sibling: content = first_sibling.text.strip() set_info['retail'] = content second_sibling = first_sibling.find_next_sibling() if second_sibling: content2 = second_sibling.text.strip() set_info['detail'] = content2 else: set_info['detail'] = "None" else: print("Not Found Retail.") sets_data.append(set_info) # Add the extracted set data to the list of all sets all_sets_data.extend(sets_data) print(f"Sets data for iteration {i + 1} extracted successfully.") # 翻页逻辑:点击Next按钮后等待新页面加载 try: next_button = WebDriverWait(driver, 10).until(EC.element_to_be_clickable((By.XPATH, "//a[contains(text(), 'Next')]"))) next_button.click() # 等待新页面数据行加载完成 WebDriverWait(driver, 10).until(EC.staleness_of(table)) time.sleep(delay_seconds) except: print("Next button not found or unclickable. Exiting loop.") break else: print("Table not found. Exiting loop.") break except Exception as e: print(f"An error occurred: {str(e)}") finally: # Close the WebDriver driver.quit() # Write all set data to a single JSON file if all_sets_data: with open('all_sets_data.json', 'w') as json_file: json.dump(all_sets_data, json_file, ensure_ascii=False, indent=4) print("All sets data extracted successfully and saved to all_sets_data.json.") else: print("No sets data extracted or saved.")
修复验证
修复后执行代码,生成的all_sets_data.json将包含完整的套装ID、名称、年份、价格等信息,不再是空对象列表。
内容的提问来源于stack exchange,提问作者BarCode
相关产品推荐
相关产品推荐

