You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页表格数据爬取求助:BeautifulSoup与Scrapy均失败

解决DWWA葡萄酒数据爬取问题

一、BeautifulSoup代码错误修正

你的错误根源是调用prettify()后将BeautifulSoup对象转为了字符串,字符串没有find_all方法。修正后的代码直接使用原始soup对象,正确遍历表格行与单元格:

import requests
from bs4 import BeautifulSoup
import pandas as pd

URL = 'https://awards.decanter.com/DWWA/2022/search/wines?competitionType=DWWA'
headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/106.0.0.0 Safari/537.36 Edg/106.0.1370.52"}

page = requests.get(URL, headers=headers)
soup = BeautifulSoup(page.content, "html.parser")

# 定位表格所有数据行(跳过表头)
table_rows = soup.select('table tbody tr')
wine_data = []

for row in table_rows:
    # 提取每行单元格的文本内容
    cells = [cell.get_text(strip=True) for cell in row.find_all('td')]
    wine_data.append(cells)

# 转为DataFrame并保存为CSV
columns = ['Producer', 'Wine', 'Vintage', 'Medal', 'Region', 'Country', 'Category']
df = pd.DataFrame(wine_data, columns=columns)
print(df.head())
df.to_csv('wines.csv', index=False)

二、Scrapy代码问题修正

你的Scrapy代码存在XPath语法错误、异步环境操作DataFrame不当、元素定位错误等问题,以下是修正后的版本:

修正核心点

  • XPath语法错误:a(@class,"dwwa-page-link") @href 应改为 //a[@class="dwwa-page-link"]/@href
  • get()方法直接返回字符串,无需再调用extract()
  • Scrapy异步环境中,避免全局操作DataFrame,通过类属性收集数据,爬取完成后统一导出
  • 页面可能存在动态渲染,若静态爬取不到数据,需结合selenium加载页面

修正后的Scrapy代码

import scrapy
import pandas as pd

class WineSpider(scrapy.Spider):
    name = 'wine_spider'
    start_urls = ["https://awards.decanter.com/DWWA/2022/search/wines?competitionType=DWWA"]
    wine_data = []

    def parse(self, response):
        # 提取当前页所有葡萄酒数据行
        table_rows = response.xpath('//table/tbody/tr')
        for row in table_rows:
            # 提取每个单元格的文本内容
            wine_item = {
                'Producer': row.xpath('./td[1]/text()').get(default='').strip(),
                'Wine': row.xpath('./td[2]/text()').get(default='').strip(),
                'Vintage': row.xpath('./td[3]/text()').get(default='').strip(),
                'Medal': row.xpath('./td[4]/text()').get(default='').strip(),
                'Region': row.xpath('./td[5]/text()').get(default='').strip(),
                'Country': row.xpath('./td[6]/text()').get(default='').strip(),
                'Category': row.xpath('./td[7]/text()').get(default='').strip()
            }
            self.wine_data.append(wine_item)
        
        # 定位下一页链接并发起请求
        next_page = response.xpath('//a[@class="dwwa-page-link" and contains(text(), "Next")]/@href').get()
        if next_page:
            yield response.follow(next_page, callback=self.parse)
        else:
            # 所有页面爬取完成后保存数据
            df = pd.DataFrame(self.wine_data)
            df.to_csv('scrapy_wines.csv', index=False)
            self.log(f"数据已保存为scrapy_wines.csv,共{len(df)}条记录")

三、额外注意事项

  • 该网站存在反爬机制,频繁请求会被封禁,建议在Scrapy中设置DOWNLOAD_DELAY = 2,或在requests代码中添加time.sleep(1)控制请求间隔
  • 若静态爬取不到数据,说明页面是JavaScript动态渲染的,可结合selenium加载页面后再解析:
    from selenium import webdriver
    import time
    from bs4 import BeautifulSoup
    
    driver = webdriver.Chrome()
    driver.get(URL)
    time.sleep(3)  # 等待页面完全加载
    soup = BeautifulSoup(driver.page_source, "html.parser")
    # 后续解析逻辑同BeautifulSoup代码
    driver.quit()
    

内容的提问来源于stack exchange,提问作者karlfranz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 11:10:24