为何Scrapy/BeautifulSoup无法爬取oxolabs.eu的Portfolio板块?
爬取oxolabs.eu Portfolio板块企业URL的问题与解决方法
问题描述
我尝试爬取网站https://oxolabs.eu/#portfolio,目标是提取Portfolio板块中的企业URL。使用Scrapy爬取时返回爬取成功但未提取到任何数据,日志如下:
2022-07-28 11:46:03 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://oxolabs.eu/?status=funded#portfolio> (referer: None)
2022-07-28 11:46:03 [scrapy.core.engine] INFO: Closing spider (finished)
使用BeautifulSoup时,返回了页面中除Portfolio板块外的所有URL。
我的BeautifulSoup脚本
from cgitb import text from re import A from bs4 import BeautifulSoup import requests url = "https://oxolabs.eu/?status=funded#portfolio" ua={'User-Agent':'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/103.0.0.0 Safari/537.36'} r = requests.get(url, headers=ua, verify=False) soup = BeautifulSoup(r.text, features="lxml") for link in soup.find_all('a'): print(link.get('href'))
我的Scrapy脚本
import scrapy class StupsbSpider(scrapy.Spider): name = 'stupsb' allowed_domains = ['oxolabs.eu/'] start_urls = ['https://oxolabs.eu/?status=funded#portfolio'] def parse(self, response): startups = response.xpath("//section[@class='oxo-section oxo-portfolio']") for startup in startups: # name = startup.xpath(".//a[@class='portfolio-entry-media-link']/@title").getall(), # industry = startup.xpath(".//div[@class='text-block-6']//text()").get(), url = startup.xpath("//section[@class='oxo-section oxo-portfolio']//@href").getall() yield{ 'url' : url, }
无法爬取的原因
- 动态内容渲染:Portfolio板块的企业数据是通过JavaScript动态加载的,
requests或Scrapy默认的HTTP请求只能获取页面初始静态HTML,动态生成的内容不会包含在响应中。 - Scrapy XPath定位瑕疵:脚本中
startup.xpath("//section[@class='oxo-section oxo-portfolio']//@href")使用了绝对路径(缺少.前缀),会从整个文档根节点重新查找,而非当前startup节点下,但这不是核心问题,核心还是动态加载。
解决方法
方法一:用浏览器渲染工具加载动态页面
使用Selenium或Playwright模拟浏览器行为,等待JavaScript执行完成后获取完整页面内容,再提取数据。
BeautifulSoup + Selenium 修改版
from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.chrome.options import Options import time url = "https://oxolabs.eu/?status=funded#portfolio" # 配置Chrome无头模式 chrome_options = Options() chrome_options.add_argument("--headless=new") chrome_options.add_argument("user-agent=Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/103.0.0.0 Safari/537.36") driver = webdriver.Chrome(options=chrome_options) driver.get(url) # 等待动态内容加载,可根据实际情况调整时长 time.sleep(3) page_source = driver.page_source driver.quit() soup = BeautifulSoup(page_source, features="lxml") portfolio_section = soup.find('section', class_='oxo-section oxo-portfolio') if portfolio_section: # 提取板块内的有效链接 for link in portfolio_section.find_all('a', href=True): href = link['href'] if href.startswith('http') and 'portfolio' not in href: print(href)
Scrapy + Selenium 修改版
import scrapy from selenium import webdriver from selenium.webdriver.chrome.options import Options from scrapy.http import HtmlResponse class StupsbSpider(scrapy.Spider): name = 'stupsb' allowed_domains = ['oxolabs.eu'] start_urls = ['https://oxolabs.eu/?status=funded#portfolio'] def __init__(self): chrome_options = Options() chrome_options.add_argument("--headless=new") chrome_options.add_argument("user-agent=Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/103.0.0.0 Safari/537.36") self.driver = webdriver.Chrome(options=chrome_options) def parse(self, response): self.driver.get(response.url) # 隐式等待5秒,确保动态内容加载完成 self.driver.implicitly_wait(5) rendered_html = self.driver.page_source # 构造Scrapy响应对象 new_response = HtmlResponse(url=response.url, body=rendered_html, encoding='utf-8') # 提取并过滤有效URL urls = new_response.xpath("//section[@class='oxo-section oxo-portfolio']//a/@href").getall() valid_urls = [url for url in urls if url.startswith('http') and 'portfolio' not in url] yield {'urls': valid_urls} def closed(self, reason): self.driver.quit()
方法二:直接调用API接口(更高效)
打开浏览器开发者工具的Network面板,刷新页面后查找加载Portfolio数据的API请求(通常是XHR或Fetch类型),直接请求该接口获取JSON格式的企业数据,从中提取URL即可,无需渲染整个页面。
内容的提问来源于stack exchange,提问作者Berci Vagyok
相关产品推荐
相关产品推荐

