You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy提取网页表格数据至CSV:请求代码实现帮助

Fixing Scrapy Table Data Extraction for KDV Yearly Values

Hey there! Let's get that table data pulled into your CSV columns properly. You've already nailed grabbing the Gemeente name, so we'll extend your existing code to target the KDV row in the table and map each year's value to your specific CSV columns.

Step-by-Step Breakdown

First, let's clarify how to target the right table cells:

  • We'll use an XPath selector to zero in on the table row where the header is "KDV"
  • Then extract each <td> element in that row, which correspond directly to the 2006, 2008, 2009, 2010, and 2011 values

Modified Parse Method

Here's the updated parse function that maps the table data to your desired CSV columns:

from collections import OrderedDict
import scrapy

class YourSpider(scrapy.Spider):
    name = 'your_spider_name'
    # Ensure this cols list includes all your target CSV columns
    cols = ['Gemeente', 'KDV_2006', 'KDV_2008', 'KDV_2009', 'KDV_2010', 'KDV_2011']

    def parse(self, response):
        item = OrderedDict()
        # Initialize all columns with 'None' as you did before
        for col in self.cols:
            item[col] = 'None'
        
        # Grab the Gemeente name (your existing working code!)
        item['Gemeente'] = response.css('h2.lijst::text').get('').strip()
        
        # Target the KDV row in the table using XPath
        kdv_row = response.xpath('//table//tr[th/text()="KDV"]/td')
        
        # Map each table cell value to the corresponding CSV column
        # Clean up whitespace and replace '-' with 'None' if needed
        if kdv_row:
            item['KDV_2006'] = kdv_row[0].xpath('./text()').get('').strip() or 'None'
            item['KDV_2008'] = kdv_row[1].xpath('./text()').get('').strip() or 'None'
            item['KDV_2009'] = kdv_row[2].xpath('./text()').get('').strip() or 'None'
            item['KDV_2010'] = kdv_row[3].xpath('./text()').get('').strip() or 'None'
            item['KDV_2011'] = kdv_row[4].xpath('./text()').get('').strip() or 'None'
        
        yield item

Key Details to Note:

  • XPath Selector: The //table//tr[th/text()="KDV"]/td selector specifically targets the row with "KDV" as its header, then fetches all data cells in that row—no guesswork about row indexes!
  • Value Cleaning: .strip() removes any extra whitespace, and or 'None' ensures empty values or '-' get replaced with your default placeholder.
  • Cols List: Double-check that your cols variable includes all the KDV_* columns you want in your CSV—this guarantees every column is initialized even if data is missing.

Bonus: Adding BSO or Totaal Data

If you need to pull BSO or Totaal values too, just replicate the pattern:

# Example for BSO row
bso_row = response.xpath('//table//tr[th/text()="BSO"]/td')
if bso_row:
    item['BSO_2006'] = bso_row[0].xpath('./text()').get('').strip() or 'None'
    # Repeat for 2008, 2009, etc.

内容的提问来源于stack exchange,提问作者Sebastiaan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 07:58:26