You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium+BeautifulSoup网页爬取报错:'NoneType'对象不可调用

问题解决:爬取Lotus's网站时触发TypeError错误

错误核心原因

你遇到的TypeError: 'NoneType' object is not callable错误,本质是BeautifulSoup方法名拼写错误:你调用了soup.findall(),但正确的方法名是soup.find_all()(注意中间的下划线)。由于findall不是BeautifulSoup的合法方法,程序返回None,将None当作函数调用就会触发这个错误。

修正后的完整代码

除了修正方法名,还优化了页面等待逻辑(确保商品元素加载完成),以下是调整后的代码:

from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

chrome_options = Options()
chrome_options.add_argument("--disable-gpu")
chrome_options.add_argument("--no-sandbox")
chrome_options.add_argument("--disable-dev-shm-usage")

service = Service(executable_path='C:/chromedriver/chromedriver.exe')
driver = webdriver.Chrome(service=service, options=chrome_options)

# 打开目标页面
driver.get('https://www.lotuss.com.my/en/category/fresh-produce?sort=relevance:DESC')

try:
    # 等待CAPTCHA验证框出现
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CSS_SELECTOR, "iframe"))
    )
    print("请在打开的浏览器窗口中手动完成CAPTCHA验证。")
finally:
    input("完成验证后按回车键继续...")
    
    # 等待商品列表元素加载完成,确保页面内容完整
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CLASS_NAME, "product-grid-item"))
    )
    
    html_text = driver.page_source
    driver.quit()

soup = BeautifulSoup(html_text, 'lxml')
# 修正方法名为find_all
grocery_items = soup.find_all('div', class_='product-grid-item')
grocery_price = soup.find_all('span', class_='sc-kHxTfl hwpbzy')

# 打印爬取结果示例
print(f"共找到{len(grocery_items)}个商品")
for item, price in zip(grocery_items, grocery_price):
    # 提取商品名称(假设h3标签存名称)
    item_name = item.find('h3').text.strip() if item.find('h3') else "未知名称"
    item_price = price.text.strip() if price else "未知价格"
    print(f"商品:{item_name} | 价格:{item_price}")
    print("---")

额外注意事项

  • 动态类名变化:网站前端的类名(如sc-kHxTfl hwpbzy)可能随框架更新改变,若后续爬取失败,需重新检查元素的CSS选择器。
  • 反爬限制:Lotus's有反爬机制,频繁请求可能触发更多验证,建议添加合理延迟,控制爬取频率。
  • 合法性:确保爬取行为符合网站robots.txt协议及使用条款,大学项目场景下建议控制爬取规模。

内容的提问来源于stack exchange,提问作者Darren Ch'ng

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 08:02:43