You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium遍历列表链接爬取详情报错:URL参数无效

问题分析与修复方案

核心错误原因

报错提示'url' must be a string,是因为遍历链接时错误传入了整个列表product_link,而非单个链接字符串;此外代码中还有多处逻辑错误,导致无法正常爬取详情页数据。

具体修复点

  • 链接调用错误:循环中driver.get(product_link)需改为driver.get(f"https://www.archify.com{product['link']}")——提取的是相对路径,必须拼接完整域名,且要取当前循环项的link值,而非整个列表。
  • 元素定位错误:By.CLASS_NAME不支持空格分隔的多类名,"text-box left-pad-25"需改用By.CSS_SELECTOR;且find_elements返回列表,不能直接调用click(),需判断存在后操作单个元素。
  • 页面元素获取错误:详情页元素不能用之前的product字典查找,需重新获取当前页面HTML并解析为BeautifulSoup对象。
  • 方法调用错误:driver.quit要加括号driver.quit(),否则不会执行浏览器关闭操作。

修改后的完整代码

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.keys import Keys
from bs4 import BeautifulSoup
import time

# 初始化浏览器
options = webdriver.ChromeOptions()
# options.add_argument('--headless')  # 如需无头模式可开启
driver = webdriver.Chrome(options=options)

url = "https://www.archify.com/id/professionals"
driver.get(url)

# 点击加载更多按钮(添加等待,避免元素未加载完成)
try:
    load_more_btn = WebDriverWait(driver, 10).until(
        EC.element_to_be_clickable((By.XPATH, "//button[text()='Load More']"))
    )
    load_more_btn.click()
except Exception as e:
    print("加载更多按钮未找到或点击失败:", e)

# 滚动页面加载更多内容
for i in range(2):
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    time.sleep(1)

# 提取列表链接
soup = BeautifulSoup(driver.page_source, 'html.parser')
product_elements = soup.find_all('div', class_='professional-box')
product_links = []
for elem in product_elements:
    text_box = elem.find('div', class_='text-box type-a')
    if text_box:
        link = text_box.find('a').get('href')
        product_links.append(link)  # 直接存储链接字符串,简化后续操作

# 遍历详情页爬取数据
product_info = []
for link in product_links:
    full_url = f"https://www.archify.com{link}"
    driver.get(full_url)
    time.sleep(2)  # 等待页面加载,或改用WebDriverWait提升效率

    # 处理展开按钮(如果存在)
    try:
        expand_btn = WebDriverWait(driver, 5).until(
            EC.element_to_be_clickable((By.CSS_SELECTOR, ".text-box.left-pad-25"))
        )
        expand_btn.click()
        time.sleep(1)
    except Exception as e:
        print(f"页面{full_url}无展开按钮或点击失败:", e)

    # 解析详情页数据
    detail_soup = BeautifulSoup(driver.page_source, 'html.parser')
    info_area = detail_soup.find('div', class_='category-list menu-left-area')
    if info_area:
        # 提取各项数据,添加异常处理避免单个字段缺失导致程序中断
        name = info_area.find('div', class_='text-box').text.strip() if info_area.find('div', class_='text-box') else "无数据"
        phone = info_area.find('div', class_='left-phone phone-number').text.strip() if info_area.find('div', class_='left-phone phone-number') else "无数据"
        website = info_area.find('div', class_='left-website phone-number').text.strip() if info_area.find('div', class_='left-website phone-number') else "无数据"
        instagram = info_area.find('div', class_='left-instagram phone-number').find('a').get('href') if info_area.find('div', class_='left-instagram phone-number') else "无数据"
        facebook = info_area.find('div', class_='left-facebook phone-number').find('a').get('href') if info_area.find('div', class_='left-facebook phone-number') else "无数据"
        whatsapp = info_area.find('div', class_='left-whatsapp phone-number').find('a').get('href') if info_area.find('div', class_='left-whatsapp phone-number') else "无数据"
        
        product_info.append({
            'Name': name,
            'Phone': phone,
            'Web': website,
            'Insta': instagram,
            'FB': facebook,
            'WA': whatsapp
        })

driver.quit()
# 可添加打印或保存数据的代码
print(product_info)

额外优化建议

  • 用WebDriverWait替代time.sleep,提升爬取稳定性和效率。
  • 对每个字段提取添加异常处理,避免因单个字段缺失导致程序崩溃。
  • 列表链接直接存储字符串而非字典,简化循环逻辑。

内容的提问来源于stack exchange,提问作者Jason

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 08:11:03