Python爬取Carousell时出现IndexError: list index out of range错误求助
问题排查与修复方案
IndexError: list index out of range 的本质是 product_card = soup.find_all('div','D_jb D_ph D_pm M_np') 返回了空列表,没有匹配到对应元素,具体问题和修复方法如下:
- 变量名语法错误
第一处是get_url函数中,定义的入参是product,拼接URL时却用了未定义的product_name变量,会直接导致生成的搜索链接无效,返回空结果页。修复方法是将url = template.format(product_name)改为url = template.format(product)。第二处是实例化Chrome的代码行外多了额外的反引号,属于语法错误,直接删掉反引号即可。 - 用动态类名定位元素不可靠
Carousell前端使用CSS Modules生成样式类名,你代码中用的D_jb D_ph D_pm M_np这类带随机后缀的类名,会随网站版本更新、页面渲染会话变化而改变,完全不具备定位稳定性,此前Shopee能用只是巧合。修复方法是换用固定属性定位,比如搜索商品卡片通用的data-testid属性,或者通过上层固定结构逐层查找,不要依赖动态类名。 - 无头模式被反爬拦截
Carousell有反爬机制,默认的无头Chrome会被识别为爬虫,返回验证页面而非正常搜索结果,页面里自然没有商品卡片。修复方法是给无头模式添加UA伪装、去掉webdriver特征标识。 - 硬等待不可靠
你用time.sleep(5)的硬等待,遇到网络波动、页面懒加载的情况,5秒可能还没渲染完商品列表。修复方法是用Selenium的显式等待,等待商品卡片元素加载完成后再解析页面。
调试建议
出现空列表的时候,先把driver.page_source打印出来保存为html文件,打开看实际返回的页面内容是什么,是链接错误、反爬验证还是元素定位错了,一目了然。
修复后参考代码
import selenium from selenium import webdriver import time from bs4 import BeautifulSoup import sys from selenium.webdriver.chrome.options import Options from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By # 生成商品搜索链接 def get_url(product): product = product.replace(' ','%20') template = 'https://www.carousell.com.my/search/{}' url = template.format(product) return url def get_all_products(card): product_image = card.find('img') product_image = product_image['src'] if product_image else '' product_name = card.find('p').text.strip() if card.find('p') else '' # 其余字段可根据实际页面固定属性调整定位逻辑,不要用动态类名 product_info = (product_image, product_name) return product_info def main(product): url = get_url(product) options = webdriver.ChromeOptions() options.add_argument('headless=new') options.add_argument('--log-level=3') # 反爬配置 options.add_argument('--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36') options.add_experimental_option('excludeSwitches', ['enable-automation']) options.add_experimental_option('useAutomationExtension', False) driver = webdriver.Chrome(executable_path='C:\\webDrivers\\chromedriver.exe',options=options) driver.execute_cdp_cmd('Page.addScriptToEvaluateOnNewDocument', { 'source': 'Object.defineProperty(navigator, "webdriver", {get: () => undefined})' }) driver.get(url) driver.maximize_window() # 显式等待商品卡片加载,最多等10秒 try: WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CSS_SELECTOR, 'div[data-testid^="listing-card"]')) ) except: print("商品列表加载超时") driver.quit() return None soup = BeautifulSoup(driver.page_source,'html.parser') product_card = soup.find_all('div', {'data-testid': lambda x: x and x.startswith('listing-card')}) # 提前判空避免索引报错 if not product_card: with open('debug.html', 'w', encoding='utf-8') as f: f.write(driver.page_source) print("未找到匹配的商品卡片,已保存页面到debug.html供排查") driver.quit() return None singleCard = product_card[0] productDetails = get_all_products(singleCard) driver.quit() return productDetails if __name__ == '__main__': if len(sys.argv) < 2: print("请传入搜索商品参数") sys.exit(1) pname = str(sys.argv[1]) scrape_data = main(pname) print(scrape_data)
内容的提问来源于stack exchange,提问作者ida khalidah
相关产品推荐
相关产品推荐

