You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

非分页站点Selenium+BS4爬取JSON为空的修复方案咨询

问题修复:Selenium+BS4爬取数据生成空JSON的解决方法

问题诊断

当前代码生成空JSON对象的核心原因:

  • 循环内重复调用driver.get(url),导致每次循环都重新加载第一页,无法实现有效翻页,且干扰页面数据渲染
  • 仅等待表格元素存在,未等待表格内的具体数据行加载完成,BeautifulSoup解析时无法定位到<tr align='left'>等有效元素
  • 缺失旧代码中价格、零售信息等关键字段的提取逻辑

修复方案(不破坏原有代码结构)

1. 修正翻页逻辑,移除循环内的重复页面加载

将driver.get(url)移到循环外部,仅在初始时加载一次页面,后续通过点击「Next」按钮实现翻页,避免重复回到第一页。

2. 优化等待条件,确保数据完全渲染

将等待表格存在改为等待表格内的数据行可见,确保页面内容完全加载后再解析:

WebDriverWait(driver, 10).until(EC.visibility_of_element_located((By.CSS_SELECTOR, "#ContentPlaceHolder1_ctlSets_GridViewSets tr[align='left']")))

3. 补全字段提取逻辑

复用旧代码中关于价格、零售信息等字段的提取逻辑,确保数据完整。

4. 翻页后增加等待,确保新页面加载完成

点击「Next」按钮后,等待新页面的数据行加载完成,再进入下一次循环解析。

完整修复代码

import json
import time
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup

# Initialize WebDriver (Safari, Chrome, Firefox, etc.)
driver = webdriver.Chrome()  # or change to webdriver.Firefox() or webdriver.Safari()

url = "https://www.brickeconomy.com/sets/year/2024"
max_iterations = 2  # Specify how many pages to fetch
delay_seconds = 2  # Delay time between each page transition (seconds)

all_sets_data = []  # List to hold all set data

try:
    # 仅初始加载一次页面
    driver.get(url)
    
    for i in range(max_iterations):
        # 等待表格内的数据行可见,确保内容完全渲染
        WebDriverWait(driver, 10).until(EC.visibility_of_element_located((By.CSS_SELECTOR, "#ContentPlaceHolder1_ctlSets_GridViewSets tr[align='left']")))

        # Process the HTML content of the page using BeautifulSoup
        soup = BeautifulSoup(driver.page_source, 'html.parser')

        sets_data = []

        # Find all rows in the table
        table = soup.find('table', id='ContentPlaceHolder1_ctlSets_GridViewSets')
        if table:
            table_rows = table.find_all('tr', align='left')

            # Extract set information from each row
            for row in table_rows:
                set_info = {}

                # Find the <h4> element containing the set name and ID
                set_name_elem = row.find('h4')
                if set_name_elem:
                    set_string = set_name_elem.text.strip()
                    set_info['id'], set_info['name'] = set_string.split(' ', 1)

                # Find <div> elements containing Year, Pieces/Minifigs, and other information
                div_elements = row.find_all('div', class_='mb-2')

                for div in div_elements:
                    label = div.find('small', class_='text-muted mr-5')
                    if label:
                        label_text = label.text.strip()

                        if label_text == 'Year':
                            set_info['year'] = div.text.replace('Year', '').strip()

                # 补全价格、零售信息提取逻辑(复用旧代码)
                td_elements = row.find_all('td', class_='ctlsets-right text-right')
                for td in td_elements:
                    div_elements = td.find_all('div')
                    for div in div_elements:
                        if "Retail" in div.text:
                            retail_price = div.text.strip()
                            price_without_retail = ' '.join(retail_price.split()[1:])
                            set_info['price'] = price_without_retail

                            first_sibling = div.find_next_sibling()
                            if first_sibling:
                                content = first_sibling.text.strip()
                                set_info['retail'] = content

                                second_sibling = first_sibling.find_next_sibling()
                                if second_sibling:
                                    content2 = second_sibling.text.strip()
                                    set_info['detail'] = content2
                                else:
                                    set_info['detail'] = "None"
                            else:
                                print("Not Found Retail.")

                sets_data.append(set_info)

            # Add the extracted set data to the list of all sets
            all_sets_data.extend(sets_data)

            print(f"Sets data for iteration {i + 1} extracted successfully.")

            # 翻页逻辑:点击Next按钮后等待新页面加载
            try:
                next_button = WebDriverWait(driver, 10).until(EC.element_to_be_clickable((By.XPATH, "//a[contains(text(), 'Next')]")))
                next_button.click()
                # 等待新页面数据行加载完成
                WebDriverWait(driver, 10).until(EC.staleness_of(table))
                time.sleep(delay_seconds)
            except:
                print("Next button not found or unclickable. Exiting loop.")
                break
        else:
            print("Table not found. Exiting loop.")
            break

except Exception as e:
    print(f"An error occurred: {str(e)}")

finally:
    # Close the WebDriver
    driver.quit()

    # Write all set data to a single JSON file
    if all_sets_data:
        with open('all_sets_data.json', 'w') as json_file:
            json.dump(all_sets_data, json_file, ensure_ascii=False, indent=4)
        print("All sets data extracted successfully and saved to all_sets_data.json.")
    else:
        print("No sets data extracted or saved.")

修复验证

修复后执行代码,生成的all_sets_data.json将包含完整的套装ID、名称、年份、价格等信息,不再是空对象列表。

内容的提问来源于stack exchange,提问作者BarCode

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 13:59:59