You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy提取表格数据至CSV:请求实现指定字段抓取代码

Scrapy Solution for Extracting KDV/BSO Table Data to CSV

Hey there! Let's get that table data extracted properly with Scrapy. I see you’ve already nailed grabbing the region name Krimpen aan den IJssel, but the table extraction is giving you trouble—let’s fix that.

The issue with your initial XPath is likely that it’s too broad; we need to target the exact rows and columns for KDV and BSO across the 2006-2011 years. Here’s a tailored Spider code snippet that should work, along with notes to adjust it to your actual HTML structure:

Step-by-Step Scrapy Code

import scrapy

class RegionTableSpider(scrapy.Spider):
    name = "region_table_spider"
    start_urls = ["YOUR_TARGET_URL_HERE"]  # Replace with your actual URL

    def parse(self, response):
        # Extract the region name (you already have this part, adjust XPath if needed)
        region = response.xpath('//*[contains(text(), "Krimpen aan den IJssel")]/text()').get().strip()

        # Target the specific table - add /tbody if your HTML uses it (common in tables)
        target_table = response.xpath('//table[@class="table paratable"]')

        # --------------------------
        # Extract KDV row data
        # --------------------------
        # Find the row containing "KDV" (adjust the XPath if KDV is in a th or nested element)
        kdv_row = target_table.xpath('.//tr[contains(., "KDV")]')
        
        # Extract values for each year - adjust the td index based on your table's column order
        # Note: XPath uses 1-based indexing!
        kdv_2006 = kdv_row.xpath('.//td[2]/text()').get(default="").strip()
        kdv_2008 = kdv_row.xpath('.//td[3]/text()').get(default="").strip()
        kdv_2010 = kdv_row.xpath('.//td[4]/text()').get(default="").strip()
        kdv_2011 = kdv_row.xpath('.//td[5]/text()').get(default="").strip()

        # --------------------------
        # Extract BSO row data
        # --------------------------
        bso_row = target_table.xpath('.//tr[contains(., "BSO")]')
        bso_2006 = bso_row.xpath('.//td[2]/text()').get(default="").strip()
        bso_2008 = bso_row.xpath('.//td[3]/text()').get(default="").strip()
        bso_2010 = bso_row.xpath('.//td[4]/text()').get(default="").strip()
        bso_2011 = bso_row.xpath('.//td[5]/text()').get(default="").strip()

        # Yield the data in a dictionary that maps directly to your CSV fields
        yield {
            "region": region,
            "KDV_2006": kdv_2006,
            "KDV_2008": kdv_2008,
            "KDV_2010": kdv_2010,
            "KDV_2011": kdv_2011,
            "BSO_2006": bso_2006,
            "BSO_2008": bso_2008,
            "BSO_2010": bso_2010,
            "BSO_2011": bso_2011
        }

Key Adjustments to Make

  • Update the start_urls: Replace with the actual URL of the page you’re scraping.
  • Adjust column indices: If 2006 isn’t in the 2nd column of the table, change the td[X] number to match your HTML’s column order.
  • Tweak row targeting: If "KDV" or "BSO" are in <th> tags instead of <td>, modify the row XPath to .//tr[contains(th/text(), "KDV")].
  • Add /tbody if needed: Many tables wrap rows in a <tbody> element—if your HTML does this, update the table XPath to //table[@class="table paratable"]/tbody.

Export to CSV

Once your spider is set up, run this command to export the data directly to a CSV file:

scrapy crawl region_table_spider -o region_data.csv

This will create a CSV with all the fields you specified, populated with the extracted table values.

内容的提问来源于stack exchange,提问作者Sebastiaan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 07:57:58