You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在BeautifulSoup中跳过指定<li>标签迭代获取缩略图链接

解决爬虫跳过无效
  • 标签的问题
  • 针对你遇到的每3个

  • 仅第一个含缩略图的情况,有两种简单的解决思路:

    思路一:按索引间隔遍历

    直接只处理索引为0、3、6...的

  • 标签,也就是步长为3的循环:

    from bs4 import BeautifulSoup
    import requests
    import openpyxl
    
    try:
        response = requests.get("https://robloxden.com/item-codes")
        soup = BeautifulSoup(response.text, 'html.parser')
        items = soup.find('ul', class_="masonry masonry--5 item-codes__container").find_all("li")
     
        # 只遍历索引为0,3,6...的元素,步长设为3
        for index in range(0, len(items), 3):
            item = items[index]
            # 提取缩略图链接
            item_link = item.find('div', class_="image-card__graphic image-card__graphic--border-bottom").img
            if item_link:  # *额外增加判断,避免意外报错*
                print(item_link['data-src'])
    
    except Exception as e:
        print(e)
    

    思路二:判断标签是否含有效内容

    遍历所有

  • ,先检查是否存在目标img标签,存在才处理:

    from bs4 import BeautifulSoup
    import requests
    import openpyxl
    
    try:
        response = requests.get("https://robloxden.com/item-codes")
        soup = BeautifulSoup(response.text, 'html.parser')
        items = soup.find('ul', class_="masonry masonry--5 item-codes__container").find_all("li")
     
        for item in items:
            # 先尝试找到目标div和img
            graphic_div = item.find('div', class_="image-card__graphic image-card__graphic--border-bottom")
            if graphic_div and graphic_div.img:
                item_link = graphic_div.img['data-src']
                print(item_link)
    
    except Exception as e:
        print(e)
    

    两种思路都能解决问题,第一种更贴合你已知的分组规律,效率更高;第二种更通用,就算页面结构小范围变动也能适配。

    内容的提问来源于stack exchange,提问作者Ammar_Fahmy

  • 相关产品推荐
    方舟 Agent Plan

    超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

    最近更新时间:2026.08.13 16:21:51