You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬虫:如何获取每个产品的首个href链接?

解决思路:只抓取每个产品的首个链接

你的代码会抓取每个产品的多个href,是因为内层循环遍历了当前产品item下所有的<a>标签(每个产品区块里通常会有图片链接、标题链接、操作按钮链接等多个a标签)。要实现每个产品只取首个href,只需修改内层的链接获取逻辑,只提取第一个符合条件的<a>标签即可,有两种简单方法:

方法1:使用find()替代find_all()

find()方法会直接返回第一个匹配到的元素,而非元素列表,刚好符合需求:

import requests
from bs4 import BeautifulSoup

baseurl = 'https://www.roco.cc/'

headers = {
    'UserAgent': 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/44.0.2403.157 Safari/537.36'
}

productlinks = []

for x in range(1,30):
    r = requests.get(
        f'https://www.roco.cc/ren/products/locomotives/steam-locomotives.html?p={x}&amp;verfuegbarkeit_status=41%2C42%2C43%2C45%2C44')

    soup = BeautifulSoup(r.content, 'lxml')
    productlist = soup.find_all('li', class_='item product product-item')
    
    for item in productlist:
        # 只获取当前产品下第一个带href的a标签
        first_link = item.find('a', href=True)
        # 增加判断避免找不到链接时报错
        if first_link:
            productlinks.append(baseurl + first_link['href'])

print(len(productlinks))

方法2:从find_all()结果中取第一个元素

如果坚持用find_all(),可以直接取返回列表的第一个元素,同时加判断避免索引越界:

# 替换内层循环部分
for item in productlist:
    all_links = item.find_all('a', href=True)
    # 确认有链接存在再取第一个
    if all_links:
        productlinks.append(baseurl + all_links[0]['href'])

两种方法都能确保每个产品只贡献一个链接,推荐用方法1,代码更简洁高效。

内容的提问来源于stack exchange,提问作者Filip Chobodicky

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 10:50:38