You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Selenium批量爬取页面?爬取10页仅获第一页结果求解

问题分析与解决方案

你的代码只拿到第一页结果,主要是几个关键的小疏漏导致的,我帮你一一拆解并修复:

1. 核心问题:未初始化Selector对象sel

你代码里直接调用sel.xpath()提取数据,但从头到尾都没给sel赋值啊!每次用Selenium打开新页面后,必须基于当前页面的源代码创建Selector实例,否则程序要么报错,要么(如果之前测试时不小心定义过sel)一直复用第一页的内容,自然只会拿到第一页的结果。

修复方法:在driver.get(url)之后,立刻加上这行代码:

sel = Selector(text=driver.page_source)

2. 爬取的数据没有写入文件

你只在开头写了CSV的表头,但循环里获取到的names、Countries、websites完全没写入文件!等于后面页面的爬取结果都丢了,最后文件里只有表头,看起来就像只爬了第一页。

你需要在循环里把数据逐行写入,用zip把三个列表对应起来(注意要确保三个列表长度一致,避免数据错位):

for name, country, website in zip(names, Countries, websites):
    f.write(f"{name.strip()}, {country.strip()}, {website.strip()}\n")

3. 可选优化:添加页面加载等待

有些网站加载速度慢,driver.get(url)之后页面还没完全渲染就提取数据,可能会拿到空结果。可以加个显式等待,确保目标元素加载完成后再提取:

from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# 在driver.get(url)之后添加
WebDriverWait(driver, 10).until(
    EC.presence_of_element_located((By.XPATH, '//*[@class="fontsubsection nomarginpadding lmargin opensans"]'))
)

修复后的完整代码

把所有问题修复后,代码应该是这样的:

# -*- coding: utf-8 -*-
from selenium import webdriver
from scrapy.selector import Selector
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

MAX_PAGE_NUM = 10
MAX_PAGE_DIG = 3
driver = webdriver.Chrome('C:\Users\zhang\Downloads\chromedriver_win32\chromedriver.exe')

with open('results.csv', 'w', encoding='utf-8') as f:
    # 修正表头,补充Country和Website字段
    f.write("Buyer, Country, Website \n")
    for i in range(1, MAX_PAGE_NUM + 1):
        page_num = (MAX_PAGE_DIG - len(str(i))) * "0" + str(i)
        url = "https://www.oilandgasnewsworldwide.com/Directory1/DREQ/Drilling_Equipment_Suppliers_?page=" + page_num
        driver.get(url)
        
        # 等待页面核心元素加载完成
        WebDriverWait(driver, 10).until(
            EC.presence_of_element_located((By.XPATH, '//*[@class="fontsubsection nomarginpadding lmargin opensans"]'))
        )
        
        # 初始化Selector,绑定当前页面源代码
        sel = Selector(text=driver.page_source)
        
        names = sel.xpath('//*[@class="fontsubsection nomarginpadding lmargin opensans"]/text()').extract()
        Countries = sel.xpath('//td[text()="Country:"]/following-sibling::td/text()').extract()
        websites = sel.xpath('//td[text()="Website:"]/following-sibling::td/a/@href').extract()
        
        # 处理并写入数据,避免空值和多余空格
        for name, country, website in zip(names, Countries, websites):
            clean_name = name.strip() if name else "N/A"
            clean_country = country.strip() if country else "N/A"
            clean_website = website.strip() if website else "N/A"
            f.write(f"{clean_name}, {clean_country}, {clean_website}\n")

driver.close()
print(len(names), len(Countries), len(websites))

额外提示

  • 打开CSV文件时指定encoding='utf-8',避免中文乱码
  • 如果三个列表长度不一致,zip会以最短的列表为准,你可以用itertools.zip_longest来处理缺失的数据
  • 频繁爬取可能触发网站反爬机制,建议添加适当延时(比如time.sleep(2),记得导入time模块)

内容的提问来源于stack exchange,提问作者Yan Zhang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 09:28:05