在KNIME的Python节点运行含Selenium的爬虫代码报错求助
解决KNIME Python Source Node中Selenium爬虫的Timeout错误
问题描述
我在Jupyter Notebook中编写的基于Selenium的网页爬虫代码运行正常,但迁移到KNIME的Python Source Node中时出现错误:
"Timeout value connect was <object object at 0x000001C6DCD29B50>, but it must be an int, float or None"
尝试了隐式等待和显式等待仍未解决,作为KNIME新手,希望得到解决建议。
错误原因与解决步骤
- 移除无效Chrome参数:Chrome浏览器不存在
--timeout命令行参数,该参数会被错误解析为无效对象,直接导致连接超时参数类型异常,需删除这一行配置。 - 统一WebDriver初始化方式:使用
Service对象管理驱动,避免同时指定executable_path和Service引发的冲突,确保KNIME环境中驱动加载正常。 - 修复异常处理漏洞:原代码中
try-except的空pass会导致products变量未定义,后续循环报错,需在异常块中给变量赋值或跳过逻辑。 - 对齐依赖环境:确认KNIME的Python环境已安装
selenium、webdriver-manager、pandas,版本尽量与Jupyter环境一致。
修改后的代码
from pandas import DataFrame from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.support.ui import Select import pandas as pd import time from webdriver_manager.chrome import ChromeDriverManager from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.chrome.options import Options # 配置Chrome选项(移除无效的--timeout参数) chrome_options = Options() # 无头模式适配无桌面环境(KNIME服务器运行时可启用) # chrome_options.add_argument("--headless=new") chrome_options.add_argument("--no-sandbox") chrome_options.add_argument("--disable-dev-shm-usage") # 使用Service初始化驱动,自动管理版本 service = Service(ChromeDriverManager().install()) driver = webdriver.Chrome(service=service, options=chrome_options) # 设置全局隐式等待 driver.implicitly_wait(10) # 访问目标网站 website = 'https://www.interpol.int/How-we-work/Notices/Red-Notices/View-Red-Notices' driver.get(website) driver.maximize_window() # 处理Cookie弹窗(显式等待确保可点击) try: cookie_btn = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.XPATH, "//i[@class='privacy-cookie-banner__icon-close']")) ) cookie_btn.click() except: pass # 初始化存储列表 name = [] ages = [] country = [] country_2 = [] testes = ['Brazil'] # 获取所有国家选项 try: select_countries = Select(WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.XPATH, '//select[@id="nationality"]')) )) for option in select_countries.options: country_2.append(option.text) # 移除第一个空占位元素 if country_2: country_2.pop(0) except: pass # 遍历测试国家 for pais in testes: try: select_countries.select_by_visible_text(pais) except: continue # 遍历年龄范围 for age in range(18, 100): try: # 填写最小年龄 min_age_input = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.XPATH, '//input[@id="ageMin"]')) ) min_age_input.clear() min_age_input.send_keys(str(age)) # 填写最大年龄 max_age_input = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.XPATH, '//input[@id="ageMax"]')) ) max_age_input.clear() max_age_input.send_keys(str(age)) # 点击搜索按钮 search_btn = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.XPATH, "//button[@id='submit' and @type='submit']")) ) search_btn.click() # 获取分页信息 pagination = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.XPATH, '//div[contains(@id, "paginationPanel")]')) ) pages = pagination.find_elements(By.XPATH, './/li') initial_page = 1 last_page = initial_page if len(pages) >= 2: try: last_page = int(pages[-2].text) except: last_page = initial_page # 分页遍历 while initial_page <= last_page: products = [] try: # 等待列表容器加载 container = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.XPATH, '//div[contains(@class, "redNoticesList__list")]')) ) # 获取所有列表项 products = WebDriverWait(container, 10).until( EC.presence_of_all_elements_located((By.XPATH, './/div[contains(@class, "redNoticesList__item notice_red")]')) ) except: # 无数据则跳过当前页 initial_page +=1 continue # 提取数据(单个字段异常时填充空值) for product in products: try: name.append(product.find_element(By.XPATH, ".//div[@class='redNoticeItem__labelText']").text) except: name.append('') try: ages.append(product.find_element(By.XPATH, './/span[@class="age"]').text) except: ages.append('') try: country.append(product.find_element(By.XPATH, './/span[@class="nationalities"]').text) except: country.append('') # 点击下一页 try: next_btn = WebDriverWait(driver, 5).until( EC.element_to_be_clickable((By.XPATH, "//a[@class='nextIndex right-arrow']")) ) next_btn.click() time.sleep(2) # 短时间等待页面切换 except: break initial_page +=1 except: continue # 统一列表长度 max_length = max(len(name), len(ages), len(country)) name += [''] * (max_length - len(name)) ages += [''] * (max_length - len(ages)) country += [''] * (max_length - len(country)) # 生成输出DataFrame df = pd.DataFrame({'Name': name, 'Age': ages, 'Country': country}) output_table = df # 关闭驱动释放资源 driver.quit()
额外提示
- 如果KNIME运行在无桌面环境(如服务器),需启用
--headless=new参数 - 尽量用显式等待替代
time.sleep(),提升爬虫稳定性 - 可在KNIME Python Source Node的日志面板查看详细报错,定位具体问题
内容的提问来源于stack exchange,提问作者Gabriel Beran Ribeiro
相关产品推荐
相关产品推荐

