网页表格数据爬取求助:BeautifulSoup与Scrapy均失败
解决DWWA葡萄酒数据爬取问题
一、BeautifulSoup代码错误修正
你的错误根源是调用prettify()后将BeautifulSoup对象转为了字符串,字符串没有find_all方法。修正后的代码直接使用原始soup对象,正确遍历表格行与单元格:
import requests from bs4 import BeautifulSoup import pandas as pd URL = 'https://awards.decanter.com/DWWA/2022/search/wines?competitionType=DWWA' headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/106.0.0.0 Safari/537.36 Edg/106.0.1370.52"} page = requests.get(URL, headers=headers) soup = BeautifulSoup(page.content, "html.parser") # 定位表格所有数据行(跳过表头) table_rows = soup.select('table tbody tr') wine_data = [] for row in table_rows: # 提取每行单元格的文本内容 cells = [cell.get_text(strip=True) for cell in row.find_all('td')] wine_data.append(cells) # 转为DataFrame并保存为CSV columns = ['Producer', 'Wine', 'Vintage', 'Medal', 'Region', 'Country', 'Category'] df = pd.DataFrame(wine_data, columns=columns) print(df.head()) df.to_csv('wines.csv', index=False)
二、Scrapy代码问题修正
你的Scrapy代码存在XPath语法错误、异步环境操作DataFrame不当、元素定位错误等问题,以下是修正后的版本:
修正核心点
- XPath语法错误:
a(@class,"dwwa-page-link") @href应改为//a[@class="dwwa-page-link"]/@href get()方法直接返回字符串,无需再调用extract()- Scrapy异步环境中,避免全局操作DataFrame,通过类属性收集数据,爬取完成后统一导出
- 页面可能存在动态渲染,若静态爬取不到数据,需结合
selenium加载页面
修正后的Scrapy代码
import scrapy import pandas as pd class WineSpider(scrapy.Spider): name = 'wine_spider' start_urls = ["https://awards.decanter.com/DWWA/2022/search/wines?competitionType=DWWA"] wine_data = [] def parse(self, response): # 提取当前页所有葡萄酒数据行 table_rows = response.xpath('//table/tbody/tr') for row in table_rows: # 提取每个单元格的文本内容 wine_item = { 'Producer': row.xpath('./td[1]/text()').get(default='').strip(), 'Wine': row.xpath('./td[2]/text()').get(default='').strip(), 'Vintage': row.xpath('./td[3]/text()').get(default='').strip(), 'Medal': row.xpath('./td[4]/text()').get(default='').strip(), 'Region': row.xpath('./td[5]/text()').get(default='').strip(), 'Country': row.xpath('./td[6]/text()').get(default='').strip(), 'Category': row.xpath('./td[7]/text()').get(default='').strip() } self.wine_data.append(wine_item) # 定位下一页链接并发起请求 next_page = response.xpath('//a[@class="dwwa-page-link" and contains(text(), "Next")]/@href').get() if next_page: yield response.follow(next_page, callback=self.parse) else: # 所有页面爬取完成后保存数据 df = pd.DataFrame(self.wine_data) df.to_csv('scrapy_wines.csv', index=False) self.log(f"数据已保存为scrapy_wines.csv,共{len(df)}条记录")
三、额外注意事项
- 该网站存在反爬机制,频繁请求会被封禁,建议在Scrapy中设置
DOWNLOAD_DELAY = 2,或在requests代码中添加time.sleep(1)控制请求间隔 - 若静态爬取不到数据,说明页面是JavaScript动态渲染的,可结合
selenium加载页面后再解析:from selenium import webdriver import time from bs4 import BeautifulSoup driver = webdriver.Chrome() driver.get(URL) time.sleep(3) # 等待页面完全加载 soup = BeautifulSoup(driver.page_source, "html.parser") # 后续解析逻辑同BeautifulSoup代码 driver.quit()
内容的提问来源于stack exchange,提问作者karlfranz
相关产品推荐
相关产品推荐

