Scrapy提取网页表格数据至CSV:请求代码实现帮助
Fixing Scrapy Table Data Extraction for KDV Yearly Values
Hey there! Let's get that table data pulled into your CSV columns properly. You've already nailed grabbing the Gemeente name, so we'll extend your existing code to target the KDV row in the table and map each year's value to your specific CSV columns.
Step-by-Step Breakdown
First, let's clarify how to target the right table cells:
- We'll use an XPath selector to zero in on the table row where the header is "KDV"
- Then extract each
<td>element in that row, which correspond directly to the 2006, 2008, 2009, 2010, and 2011 values
Modified Parse Method
Here's the updated parse function that maps the table data to your desired CSV columns:
from collections import OrderedDict import scrapy class YourSpider(scrapy.Spider): name = 'your_spider_name' # Ensure this cols list includes all your target CSV columns cols = ['Gemeente', 'KDV_2006', 'KDV_2008', 'KDV_2009', 'KDV_2010', 'KDV_2011'] def parse(self, response): item = OrderedDict() # Initialize all columns with 'None' as you did before for col in self.cols: item[col] = 'None' # Grab the Gemeente name (your existing working code!) item['Gemeente'] = response.css('h2.lijst::text').get('').strip() # Target the KDV row in the table using XPath kdv_row = response.xpath('//table//tr[th/text()="KDV"]/td') # Map each table cell value to the corresponding CSV column # Clean up whitespace and replace '-' with 'None' if needed if kdv_row: item['KDV_2006'] = kdv_row[0].xpath('./text()').get('').strip() or 'None' item['KDV_2008'] = kdv_row[1].xpath('./text()').get('').strip() or 'None' item['KDV_2009'] = kdv_row[2].xpath('./text()').get('').strip() or 'None' item['KDV_2010'] = kdv_row[3].xpath('./text()').get('').strip() or 'None' item['KDV_2011'] = kdv_row[4].xpath('./text()').get('').strip() or 'None' yield item
Key Details to Note:
- XPath Selector: The
//table//tr[th/text()="KDV"]/tdselector specifically targets the row with "KDV" as its header, then fetches all data cells in that row—no guesswork about row indexes! - Value Cleaning:
.strip()removes any extra whitespace, andor 'None'ensures empty values or '-' get replaced with your default placeholder. - Cols List: Double-check that your
colsvariable includes all the KDV_* columns you want in your CSV—this guarantees every column is initialized even if data is missing.
Bonus: Adding BSO or Totaal Data
If you need to pull BSO or Totaal values too, just replicate the pattern:
# Example for BSO row bso_row = response.xpath('//table//tr[th/text()="BSO"]/td') if bso_row: item['BSO_2006'] = bso_row[0].xpath('./text()').get('').strip() or 'None' # Repeat for 2008, 2009, etc.
内容的提问来源于stack exchange,提问作者Sebastiaan
相关产品推荐
相关产品推荐

