Scrapy提取表格数据至CSV:请求实现指定字段抓取代码
Hey there! Let's get that table data extracted properly with Scrapy. I see you’ve already nailed grabbing the region name Krimpen aan den IJssel, but the table extraction is giving you trouble—let’s fix that.
The issue with your initial XPath is likely that it’s too broad; we need to target the exact rows and columns for KDV and BSO across the 2006-2011 years. Here’s a tailored Spider code snippet that should work, along with notes to adjust it to your actual HTML structure:
Step-by-Step Scrapy Code
import scrapy class RegionTableSpider(scrapy.Spider): name = "region_table_spider" start_urls = ["YOUR_TARGET_URL_HERE"] # Replace with your actual URL def parse(self, response): # Extract the region name (you already have this part, adjust XPath if needed) region = response.xpath('//*[contains(text(), "Krimpen aan den IJssel")]/text()').get().strip() # Target the specific table - add /tbody if your HTML uses it (common in tables) target_table = response.xpath('//table[@class="table paratable"]') # -------------------------- # Extract KDV row data # -------------------------- # Find the row containing "KDV" (adjust the XPath if KDV is in a th or nested element) kdv_row = target_table.xpath('.//tr[contains(., "KDV")]') # Extract values for each year - adjust the td index based on your table's column order # Note: XPath uses 1-based indexing! kdv_2006 = kdv_row.xpath('.//td[2]/text()').get(default="").strip() kdv_2008 = kdv_row.xpath('.//td[3]/text()').get(default="").strip() kdv_2010 = kdv_row.xpath('.//td[4]/text()').get(default="").strip() kdv_2011 = kdv_row.xpath('.//td[5]/text()').get(default="").strip() # -------------------------- # Extract BSO row data # -------------------------- bso_row = target_table.xpath('.//tr[contains(., "BSO")]') bso_2006 = bso_row.xpath('.//td[2]/text()').get(default="").strip() bso_2008 = bso_row.xpath('.//td[3]/text()').get(default="").strip() bso_2010 = bso_row.xpath('.//td[4]/text()').get(default="").strip() bso_2011 = bso_row.xpath('.//td[5]/text()').get(default="").strip() # Yield the data in a dictionary that maps directly to your CSV fields yield { "region": region, "KDV_2006": kdv_2006, "KDV_2008": kdv_2008, "KDV_2010": kdv_2010, "KDV_2011": kdv_2011, "BSO_2006": bso_2006, "BSO_2008": bso_2008, "BSO_2010": bso_2010, "BSO_2011": bso_2011 }
Key Adjustments to Make
- Update the
start_urls: Replace with the actual URL of the page you’re scraping. - Adjust column indices: If 2006 isn’t in the 2nd column of the table, change the
td[X]number to match your HTML’s column order. - Tweak row targeting: If "KDV" or "BSO" are in
<th>tags instead of<td>, modify the row XPath to.//tr[contains(th/text(), "KDV")]. - Add
/tbodyif needed: Many tables wrap rows in a<tbody>element—if your HTML does this, update the table XPath to//table[@class="table paratable"]/tbody.
Export to CSV
Once your spider is set up, run this command to export the data directly to a CSV file:
scrapy crawl region_table_spider -o region_data.csv
This will create a CSV with all the fields you specified, populated with the extracted table values.
内容的提问来源于stack exchange,提问作者Sebastiaan

