Selenium遍历列表链接爬取详情报错:URL参数无效
问题分析与修复方案
核心错误原因
报错提示'url' must be a string,是因为遍历链接时错误传入了整个列表product_link,而非单个链接字符串;此外代码中还有多处逻辑错误,导致无法正常爬取详情页数据。
具体修复点
- 链接调用错误:循环中
driver.get(product_link)需改为driver.get(f"https://www.archify.com{product['link']}")——提取的是相对路径,必须拼接完整域名,且要取当前循环项的link值,而非整个列表。 - 元素定位错误:
By.CLASS_NAME不支持空格分隔的多类名,"text-box left-pad-25"需改用By.CSS_SELECTOR;且find_elements返回列表,不能直接调用click(),需判断存在后操作单个元素。 - 页面元素获取错误:详情页元素不能用之前的
product字典查找,需重新获取当前页面HTML并解析为BeautifulSoup对象。 - 方法调用错误:
driver.quit要加括号driver.quit(),否则不会执行浏览器关闭操作。
修改后的完整代码
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.keys import Keys from bs4 import BeautifulSoup import time # 初始化浏览器 options = webdriver.ChromeOptions() # options.add_argument('--headless') # 如需无头模式可开启 driver = webdriver.Chrome(options=options) url = "https://www.archify.com/id/professionals" driver.get(url) # 点击加载更多按钮(添加等待,避免元素未加载完成) try: load_more_btn = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.XPATH, "//button[text()='Load More']")) ) load_more_btn.click() except Exception as e: print("加载更多按钮未找到或点击失败:", e) # 滚动页面加载更多内容 for i in range(2): driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") time.sleep(1) # 提取列表链接 soup = BeautifulSoup(driver.page_source, 'html.parser') product_elements = soup.find_all('div', class_='professional-box') product_links = [] for elem in product_elements: text_box = elem.find('div', class_='text-box type-a') if text_box: link = text_box.find('a').get('href') product_links.append(link) # 直接存储链接字符串,简化后续操作 # 遍历详情页爬取数据 product_info = [] for link in product_links: full_url = f"https://www.archify.com{link}" driver.get(full_url) time.sleep(2) # 等待页面加载,或改用WebDriverWait提升效率 # 处理展开按钮(如果存在) try: expand_btn = WebDriverWait(driver, 5).until( EC.element_to_be_clickable((By.CSS_SELECTOR, ".text-box.left-pad-25")) ) expand_btn.click() time.sleep(1) except Exception as e: print(f"页面{full_url}无展开按钮或点击失败:", e) # 解析详情页数据 detail_soup = BeautifulSoup(driver.page_source, 'html.parser') info_area = detail_soup.find('div', class_='category-list menu-left-area') if info_area: # 提取各项数据,添加异常处理避免单个字段缺失导致程序中断 name = info_area.find('div', class_='text-box').text.strip() if info_area.find('div', class_='text-box') else "无数据" phone = info_area.find('div', class_='left-phone phone-number').text.strip() if info_area.find('div', class_='left-phone phone-number') else "无数据" website = info_area.find('div', class_='left-website phone-number').text.strip() if info_area.find('div', class_='left-website phone-number') else "无数据" instagram = info_area.find('div', class_='left-instagram phone-number').find('a').get('href') if info_area.find('div', class_='left-instagram phone-number') else "无数据" facebook = info_area.find('div', class_='left-facebook phone-number').find('a').get('href') if info_area.find('div', class_='left-facebook phone-number') else "无数据" whatsapp = info_area.find('div', class_='left-whatsapp phone-number').find('a').get('href') if info_area.find('div', class_='left-whatsapp phone-number') else "无数据" product_info.append({ 'Name': name, 'Phone': phone, 'Web': website, 'Insta': instagram, 'FB': facebook, 'WA': whatsapp }) driver.quit() # 可添加打印或保存数据的代码 print(product_info)
额外优化建议
- 用
WebDriverWait替代time.sleep,提升爬取稳定性和效率。 - 对每个字段提取添加异常处理,避免因单个字段缺失导致程序崩溃。
- 列表链接直接存储字符串而非字典,简化循环逻辑。
内容的提问来源于stack exchange,提问作者Jason
相关产品推荐
相关产品推荐

