You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在KNIME的Python节点运行含Selenium的爬虫代码报错求助

解决KNIME Python Source Node中Selenium爬虫的Timeout错误

问题描述

我在Jupyter Notebook中编写的基于Selenium的网页爬虫代码运行正常,但迁移到KNIME的Python Source Node中时出现错误:

"Timeout value connect was <object object at 0x000001C6DCD29B50>, but it must be an int, float or None"

尝试了隐式等待和显式等待仍未解决,作为KNIME新手,希望得到解决建议。

错误原因与解决步骤

  1. 移除无效Chrome参数:Chrome浏览器不存在--timeout命令行参数,该参数会被错误解析为无效对象,直接导致连接超时参数类型异常,需删除这一行配置。
  2. 统一WebDriver初始化方式:使用Service对象管理驱动,避免同时指定executable_path和Service引发的冲突,确保KNIME环境中驱动加载正常。
  3. 修复异常处理漏洞:原代码中try-except的空pass会导致products变量未定义,后续循环报错,需在异常块中给变量赋值或跳过逻辑。
  4. 对齐依赖环境:确认KNIME的Python环境已安装selenium、webdriver-manager、pandas,版本尽量与Jupyter环境一致。

修改后的代码

from pandas import DataFrame
from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.support.ui import Select
import pandas as pd
import time
from webdriver_manager.chrome import ChromeDriverManager
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.chrome.options import Options

# 配置Chrome选项(移除无效的--timeout参数)
chrome_options = Options()
# 无头模式适配无桌面环境(KNIME服务器运行时可启用)
# chrome_options.add_argument("--headless=new")
chrome_options.add_argument("--no-sandbox")
chrome_options.add_argument("--disable-dev-shm-usage")

# 使用Service初始化驱动,自动管理版本
service = Service(ChromeDriverManager().install())
driver = webdriver.Chrome(service=service, options=chrome_options)

# 设置全局隐式等待
driver.implicitly_wait(10)

# 访问目标网站
website = 'https://www.interpol.int/How-we-work/Notices/Red-Notices/View-Red-Notices'
driver.get(website)
driver.maximize_window()

# 处理Cookie弹窗(显式等待确保可点击)
try:
    cookie_btn = WebDriverWait(driver, 10).until(
        EC.element_to_be_clickable((By.XPATH, "//i[@class='privacy-cookie-banner__icon-close']"))
    )
    cookie_btn.click()
except:
    pass

# 初始化存储列表
name = []
ages = []
country = []
country_2 = []
testes = ['Brazil']

# 获取所有国家选项
try:
    select_countries = Select(WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.XPATH, '//select[@id="nationality"]'))
    ))
    for option in select_countries.options:
        country_2.append(option.text)
    # 移除第一个空占位元素
    if country_2:
        country_2.pop(0)
except:
    pass

# 遍历测试国家
for pais in testes:
    try:
        select_countries.select_by_visible_text(pais)
    except:
        continue

    # 遍历年龄范围
    for age in range(18, 100):
        try:
            # 填写最小年龄
            min_age_input = WebDriverWait(driver, 10).until(
                EC.presence_of_element_located((By.XPATH, '//input[@id="ageMin"]'))
            )
            min_age_input.clear()
            min_age_input.send_keys(str(age))

            # 填写最大年龄
            max_age_input = WebDriverWait(driver, 10).until(
                EC.presence_of_element_located((By.XPATH, '//input[@id="ageMax"]'))
            )
            max_age_input.clear()
            max_age_input.send_keys(str(age))

            # 点击搜索按钮
            search_btn = WebDriverWait(driver, 10).until(
                EC.element_to_be_clickable((By.XPATH, "//button[@id='submit' and @type='submit']"))
            )
            search_btn.click()

            # 获取分页信息
            pagination = WebDriverWait(driver, 10).until(
                EC.presence_of_element_located((By.XPATH, '//div[contains(@id, "paginationPanel")]'))
            )
            pages = pagination.find_elements(By.XPATH, './/li')
            initial_page = 1
            last_page = initial_page

            if len(pages) >= 2:
                try:
                    last_page = int(pages[-2].text)
                except:
                    last_page = initial_page

            # 分页遍历
            while initial_page <= last_page:
                products = []
                try:
                    # 等待列表容器加载
                    container = WebDriverWait(driver, 10).until(
                        EC.presence_of_element_located((By.XPATH, '//div[contains(@class, "redNoticesList__list")]'))
                    )
                    # 获取所有列表项
                    products = WebDriverWait(container, 10).until(
                        EC.presence_of_all_elements_located((By.XPATH, './/div[contains(@class, "redNoticesList__item notice_red")]'))
                    )
                except:
                    # 无数据则跳过当前页
                    initial_page +=1
                    continue

                # 提取数据(单个字段异常时填充空值)
                for product in products:
                    try:
                        name.append(product.find_element(By.XPATH, ".//div[@class='redNoticeItem__labelText']").text)
                    except:
                        name.append('')
                    try:
                        ages.append(product.find_element(By.XPATH, './/span[@class="age"]').text)
                    except:
                        ages.append('')
                    try:
                        country.append(product.find_element(By.XPATH, './/span[@class="nationalities"]').text)
                    except:
                        country.append('')

                # 点击下一页
                try:
                    next_btn = WebDriverWait(driver, 5).until(
                        EC.element_to_be_clickable((By.XPATH, "//a[@class='nextIndex right-arrow']"))
                    )
                    next_btn.click()
                    time.sleep(2)  # 短时间等待页面切换
                except:
                    break

                initial_page +=1

        except:
            continue

# 统一列表长度
max_length = max(len(name), len(ages), len(country))
name += [''] * (max_length - len(name))
ages += [''] * (max_length - len(ages))
country += [''] * (max_length - len(country))

# 生成输出DataFrame
df = pd.DataFrame({'Name': name, 'Age': ages, 'Country': country})
output_table = df

# 关闭驱动释放资源
driver.quit()

额外提示

  • 如果KNIME运行在无桌面环境(如服务器),需启用--headless=new参数
  • 尽量用显式等待替代time.sleep(),提升爬虫稳定性
  • 可在KNIME Python Source Node的日志面板查看详细报错,定位具体问题

内容的提问来源于stack exchange,提问作者Gabriel Beran Ribeiro

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 00:30:56