如何使用Scrapy存储网页表格?新手爬取指定表格遇存储难题
Hey there! Since you're new to Scrapy and struggling to extract that table with ID grdTableView_DXMainTable from the World Bank Databank page, let's work through this together. Here's a complete, actionable solution:
1. First, Complete Your Basic Spider Code
Your start_requests was incomplete, so let's fill that out and add a parse method to target the table:
import scrapy class EducationSpider(scrapy.Spider): name = "education" allowed_domains = ["databank.worldbank.org"] # Use start_urls directly instead of start_requests for simplicity (it's equivalent) start_urls = [ 'http://databank.worldbank.org/data/reports.aspx?source=2&series=SE.PRM.NENR&country=' ] def parse(self, response): # Target the table using its ID with CSS selector target_table = response.css('#grdTableView_DXMainTable') if not target_table: self.logger.warning("Oops! Couldn't find the target table. This might be because the page loads content dynamically.") return # Extract the full HTML of the table table_html = target_table.get() # Save the table to an HTML file (you can also save as CSV/JSON later if needed) with open('worldbank_primary_education_table.html', 'w', encoding='utf-8') as output_file: output_file.write(table_html) self.logger.info("Success! The table has been saved to your file.")
2. What If The Table Doesn't Show Up?
The World Bank Databank uses DevExpress UI components, which often load tables via AJAX (dynamic content) instead of including them in the initial page HTML. If the code above doesn't find the table, try one of these approaches:
Option A: Fetch Data Directly From The AJAX API
Open your browser's DevTools (F12), go to the Network tab, refresh the page, and look for API requests (usually with json in the name). You'll likely find a request that returns the table data in JSON format — this is way more efficient than parsing HTML. Just update your spider to call that API endpoint instead!
Option B: Use Selenium To Render Dynamic Pages
If you don't want to hunt for the API, use Selenium to let a browser render the page fully before scraping:
- Install the required package:
pip install scrapy-selenium - Add these settings to your
settings.py:DOWNLOADER_MIDDLEWARES = { 'scrapy_selenium.SeleniumMiddleware': 800 } SELENIUM_DRIVER_NAME = 'chrome' SELENIUM_DRIVER_EXECUTABLE_PATH = '/path/to/your/chromedriver' # Update this path! SELENIUM_DRIVER_ARGUMENTS = ['--headless=new'] # Run in headless mode (no browser window) - Update your spider to use
SeleniumRequest:import scrapy from scrapy_selenium import SeleniumRequest class EducationSpider(scrapy.Spider): name = "education" allowed_domains = ["databank.worldbank.org"] def start_requests(self): yield SeleniumRequest( url='http://databank.worldbank.org/data/reports.aspx?source=2&series=SE.PRM.NENR&country=', wait_time=5, # Give the page time to load wait_until='element_to_be_clickable', wait_for='#grdTableView_DXMainTable', # Wait until the table is visible callback=self.parse ) def parse(self, response): target_table = response.css('#grdTableView_DXMainTable') table_html = target_table.get() with open('worldbank_primary_education_table.html', 'w', encoding='utf-8') as output_file: output_file.write(table_html) self.logger.info("Dynamic table saved successfully!")
3. Bonus: Extract Specific Table Data (Not Just HTML)
If you don't want the full HTML and instead want to pull out rows/columns, add this to your parse method:
# Inside parse(), after getting target_table rows = target_table.css('tr') for row in rows: # Extract text from all header and data cells cells = row.css('th::text, td::text').getall() # Clean up whitespace cleaned_cells = [cell.strip() for cell in cells if cell.strip()] self.logger.info(f"Row data: {cleaned_cells}") # You can save this to a CSV/JSON file using Scrapy's Item pipelines too!
Just a quick reminder: Always respect the site's robots.txt and don't spam requests — you don't want to get your IP blocked.
内容的提问来源于stack exchange,提问作者Roberta Gimenez

