Selenium+BeautifulSoup网页爬取报错:'NoneType'对象不可调用
问题解决:爬取Lotus's网站时触发TypeError错误
错误核心原因
你遇到的TypeError: 'NoneType' object is not callable错误,本质是BeautifulSoup方法名拼写错误:你调用了soup.findall(),但正确的方法名是soup.find_all()(注意中间的下划线)。由于findall不是BeautifulSoup的合法方法,程序返回None,将None当作函数调用就会触发这个错误。
修正后的完整代码
除了修正方法名,还优化了页面等待逻辑(确保商品元素加载完成),以下是调整后的代码:
from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.common.by import By from selenium.webdriver.chrome.options import Options from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC chrome_options = Options() chrome_options.add_argument("--disable-gpu") chrome_options.add_argument("--no-sandbox") chrome_options.add_argument("--disable-dev-shm-usage") service = Service(executable_path='C:/chromedriver/chromedriver.exe') driver = webdriver.Chrome(service=service, options=chrome_options) # 打开目标页面 driver.get('https://www.lotuss.com.my/en/category/fresh-produce?sort=relevance:DESC') try: # 等待CAPTCHA验证框出现 WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CSS_SELECTOR, "iframe")) ) print("请在打开的浏览器窗口中手动完成CAPTCHA验证。") finally: input("完成验证后按回车键继续...") # 等待商品列表元素加载完成,确保页面内容完整 WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, "product-grid-item")) ) html_text = driver.page_source driver.quit() soup = BeautifulSoup(html_text, 'lxml') # 修正方法名为find_all grocery_items = soup.find_all('div', class_='product-grid-item') grocery_price = soup.find_all('span', class_='sc-kHxTfl hwpbzy') # 打印爬取结果示例 print(f"共找到{len(grocery_items)}个商品") for item, price in zip(grocery_items, grocery_price): # 提取商品名称(假设h3标签存名称) item_name = item.find('h3').text.strip() if item.find('h3') else "未知名称" item_price = price.text.strip() if price else "未知价格" print(f"商品:{item_name} | 价格:{item_price}") print("---")
额外注意事项
- 动态类名变化:网站前端的类名(如
sc-kHxTfl hwpbzy)可能随框架更新改变,若后续爬取失败,需重新检查元素的CSS选择器。 - 反爬限制:Lotus's有反爬机制,频繁请求可能触发更多验证,建议添加合理延迟,控制爬取频率。
- 合法性:确保爬取行为符合网站
robots.txt协议及使用条款,大学项目场景下建议控制爬取规模。
内容的提问来源于stack exchange,提问作者Darren Ch'ng
相关产品推荐
相关产品推荐

