You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何Scrapy/BeautifulSoup无法爬取oxolabs.eu的Portfolio板块?

爬取oxolabs.eu Portfolio板块企业URL的问题与解决方法

问题描述

我尝试爬取网站https://oxolabs.eu/#portfolio,目标是提取Portfolio板块中的企业URL。使用Scrapy爬取时返回爬取成功但未提取到任何数据,日志如下:

2022-07-28 11:46:03 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://oxolabs.eu/?status=funded#portfolio> (referer: None)
2022-07-28 11:46:03 [scrapy.core.engine] INFO: Closing spider (finished)

使用BeautifulSoup时,返回了页面中除Portfolio板块外的所有URL。

我的BeautifulSoup脚本

from cgitb import text
from re import A
from bs4 import BeautifulSoup
import requests

url = "https://oxolabs.eu/?status=funded#portfolio"
ua={'User-Agent':'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/103.0.0.0 Safari/537.36'}
r = requests.get(url, headers=ua, verify=False)
soup = BeautifulSoup(r.text, features="lxml")

for link in soup.find_all('a'):
    print(link.get('href'))

我的Scrapy脚本

import scrapy


class StupsbSpider(scrapy.Spider):
    name = 'stupsb'
    allowed_domains = ['oxolabs.eu/']
    start_urls = ['https://oxolabs.eu/?status=funded#portfolio']

    def parse(self, response):
        startups = response.xpath("//section[@class='oxo-section oxo-portfolio']")
        for startup in startups:
            # name = startup.xpath(".//a[@class='portfolio-entry-media-link']/@title").getall(),
            # industry = startup.xpath(".//div[@class='text-block-6']//text()").get(),
            url = startup.xpath("//section[@class='oxo-section oxo-portfolio']//@href").getall()
            yield{
                'url' : url,
            }

无法爬取的原因

  • 动态内容渲染:Portfolio板块的企业数据是通过JavaScript动态加载的,requests或Scrapy默认的HTTP请求只能获取页面初始静态HTML,动态生成的内容不会包含在响应中。
  • Scrapy XPath定位瑕疵:脚本中startup.xpath("//section[@class='oxo-section oxo-portfolio']//@href")使用了绝对路径(缺少.前缀),会从整个文档根节点重新查找,而非当前startup节点下,但这不是核心问题,核心还是动态加载。

解决方法

方法一:用浏览器渲染工具加载动态页面

使用Selenium或Playwright模拟浏览器行为,等待JavaScript执行完成后获取完整页面内容,再提取数据。

BeautifulSoup + Selenium 修改版

from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
import time

url = "https://oxolabs.eu/?status=funded#portfolio"
# 配置Chrome无头模式
chrome_options = Options()
chrome_options.add_argument("--headless=new")
chrome_options.add_argument("user-agent=Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/103.0.0.0 Safari/537.36")

driver = webdriver.Chrome(options=chrome_options)
driver.get(url)
# 等待动态内容加载,可根据实际情况调整时长
time.sleep(3)
page_source = driver.page_source
driver.quit()

soup = BeautifulSoup(page_source, features="lxml")
portfolio_section = soup.find('section', class_='oxo-section oxo-portfolio')
if portfolio_section:
    # 提取板块内的有效链接
    for link in portfolio_section.find_all('a', href=True):
        href = link['href']
        if href.startswith('http') and 'portfolio' not in href:
            print(href)

Scrapy + Selenium 修改版

import scrapy
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from scrapy.http import HtmlResponse


class StupsbSpider(scrapy.Spider):
    name = 'stupsb'
    allowed_domains = ['oxolabs.eu']
    start_urls = ['https://oxolabs.eu/?status=funded#portfolio']

    def __init__(self):
        chrome_options = Options()
        chrome_options.add_argument("--headless=new")
        chrome_options.add_argument("user-agent=Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/103.0.0.0 Safari/537.36")
        self.driver = webdriver.Chrome(options=chrome_options)

    def parse(self, response):
        self.driver.get(response.url)
        # 隐式等待5秒,确保动态内容加载完成
        self.driver.implicitly_wait(5)
        rendered_html = self.driver.page_source
        # 构造Scrapy响应对象
        new_response = HtmlResponse(url=response.url, body=rendered_html, encoding='utf-8')
        
        # 提取并过滤有效URL
        urls = new_response.xpath("//section[@class='oxo-section oxo-portfolio']//a/@href").getall()
        valid_urls = [url for url in urls if url.startswith('http') and 'portfolio' not in url]
        yield {'urls': valid_urls}

    def closed(self, reason):
        self.driver.quit()

方法二:直接调用API接口(更高效)

打开浏览器开发者工具的Network面板,刷新页面后查找加载Portfolio数据的API请求(通常是XHR或Fetch类型),直接请求该接口获取JSON格式的企业数据,从中提取URL即可,无需渲染整个页面。


内容的提问来源于stack exchange,提问作者Berci Vagyok

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 10:27:28