You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Scrapy存储网页表格?新手爬取指定表格遇存储难题

Hey there! Since you're new to Scrapy and struggling to extract that table with ID grdTableView_DXMainTable from the World Bank Databank page, let's work through this together. Here's a complete, actionable solution:

Fixing Your Scrapy Table Extraction

1. First, Complete Your Basic Spider Code

Your start_requests was incomplete, so let's fill that out and add a parse method to target the table:

import scrapy

class EducationSpider(scrapy.Spider):
    name = "education"
    allowed_domains = ["databank.worldbank.org"]
    # Use start_urls directly instead of start_requests for simplicity (it's equivalent)
    start_urls = [
        'http://databank.worldbank.org/data/reports.aspx?source=2&series=SE.PRM.NENR&country='
    ]

    def parse(self, response):
        # Target the table using its ID with CSS selector
        target_table = response.css('#grdTableView_DXMainTable')

        if not target_table:
            self.logger.warning("Oops! Couldn't find the target table. This might be because the page loads content dynamically.")
            return

        # Extract the full HTML of the table
        table_html = target_table.get()

        # Save the table to an HTML file (you can also save as CSV/JSON later if needed)
        with open('worldbank_primary_education_table.html', 'w', encoding='utf-8') as output_file:
            output_file.write(table_html)

        self.logger.info("Success! The table has been saved to your file.")

2. What If The Table Doesn't Show Up?

The World Bank Databank uses DevExpress UI components, which often load tables via AJAX (dynamic content) instead of including them in the initial page HTML. If the code above doesn't find the table, try one of these approaches:

Option A: Fetch Data Directly From The AJAX API

Open your browser's DevTools (F12), go to the Network tab, refresh the page, and look for API requests (usually with json in the name). You'll likely find a request that returns the table data in JSON format — this is way more efficient than parsing HTML. Just update your spider to call that API endpoint instead!

Option B: Use Selenium To Render Dynamic Pages

If you don't want to hunt for the API, use Selenium to let a browser render the page fully before scraping:

  1. Install the required package:
    pip install scrapy-selenium
    
  2. Add these settings to your settings.py:
    DOWNLOADER_MIDDLEWARES = {
        'scrapy_selenium.SeleniumMiddleware': 800
    }
    SELENIUM_DRIVER_NAME = 'chrome'
    SELENIUM_DRIVER_EXECUTABLE_PATH = '/path/to/your/chromedriver' # Update this path!
    SELENIUM_DRIVER_ARGUMENTS = ['--headless=new'] # Run in headless mode (no browser window)
    
  3. Update your spider to use SeleniumRequest:
    import scrapy
    from scrapy_selenium import SeleniumRequest
    
    class EducationSpider(scrapy.Spider):
        name = "education"
        allowed_domains = ["databank.worldbank.org"]
    
        def start_requests(self):
            yield SeleniumRequest(
                url='http://databank.worldbank.org/data/reports.aspx?source=2&series=SE.PRM.NENR&country=',
                wait_time=5, # Give the page time to load
                wait_until='element_to_be_clickable',
                wait_for='#grdTableView_DXMainTable', # Wait until the table is visible
                callback=self.parse
            )
    
        def parse(self, response):
            target_table = response.css('#grdTableView_DXMainTable')
            table_html = target_table.get()
    
            with open('worldbank_primary_education_table.html', 'w', encoding='utf-8') as output_file:
                output_file.write(table_html)
    
            self.logger.info("Dynamic table saved successfully!")
    

3. Bonus: Extract Specific Table Data (Not Just HTML)

If you don't want the full HTML and instead want to pull out rows/columns, add this to your parse method:

# Inside parse(), after getting target_table
rows = target_table.css('tr')
for row in rows:
    # Extract text from all header and data cells
    cells = row.css('th::text, td::text').getall()
    # Clean up whitespace
    cleaned_cells = [cell.strip() for cell in cells if cell.strip()]
    self.logger.info(f"Row data: {cleaned_cells}")
    # You can save this to a CSV/JSON file using Scrapy's Item pipelines too!

Just a quick reminder: Always respect the site's robots.txt and don't spam requests — you don't want to get your IP blocked.

内容的提问来源于stack exchange,提问作者Roberta Gimenez

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:44:33