You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

LinkedIn Ads Library爬取问题:无法获取图片src属性

解决LinkedIn Ads Library图片及广告内容爬取问题

问题分析

LinkedIn Ads Library的部分内容(包括图片链接)依赖JavaScript动态渲染,直接用requests获取的静态HTML中,图片标签常缺失src属性,真实链接会存储在data-*类属性中;同时,未携带合法请求头的请求易被LinkedIn拦截,导致返回内容不全。

改进现有BeautifulSoup方案

1. 补充请求头模拟浏览器行为

添加真实浏览器的User-Agent,避免被LinkedIn的反爬机制拦截,提升内容获取的完整性。

2. 修正图片属性提取逻辑

实际页面中,广告图片的真实链接通常存储在data-delayed-url属性中,调整提取逻辑优先读取该属性,再降级到src。

改进后的代码:

import requests
from bs4 import BeautifulSoup

def LinkedInAdScraper(company: str):
    url = f"https://www.linkedin.com/ad-library/search?accountOwner={company}&dateOption=last-30-days"
    
    # 模拟浏览器请求头
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
    }

    response = requests.get(url, headers=headers)

    if response.status_code == 200:
        soup = BeautifulSoup(response.content, 'html.parser')
        ad_previews = soup.find_all('div', class_='base-ad-preview-card')

        for ad_preview in ad_previews:
            # 提取图片链接:优先读取data-delayed-url,无则 fallback 到src
            image_tag = ad_preview.select_one('.ad-preview__dynamic-dimensions-image')
            image_link = image_tag.get('data-delayed-url') or image_tag.get('src') if image_tag else None
            
            # 提取标题和内容
            title = ad_preview.select_one('.text-md').text.strip() if ad_preview.select_one('.text-md') else None
            content = ad_preview.select_one('.commentary__content').text.strip() if ad_preview.select_one('.commentary__content') else None

            print("Image Link:", image_link)
            print("Title:", title)
            print("Content:", content)
            print("-" * 50)
    else:
        print(f"请求失败,状态码: {response.status_code}")

LinkedInAdScraper("SalesForce")

动态渲染内容爬取方案(Selenium)

若上述方法仍无法获取完整内容(比如部分广告为异步加载),可使用Selenium模拟浏览器渲染页面,确保所有动态内容加载完成后再提取:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time
from bs4 import BeautifulSoup

def LinkedInAdScraperSelenium(company: str):
    url = f"https://www.linkedin.com/ad-library/search?accountOwner={company}&dateOption=last-30-days"
    
    # 初始化Chrome浏览器(需提前安装ChromeDriver并配置环境变量)
    options = webdriver.ChromeOptions()
    options.add_argument("--headless=new")  # 无头模式,不显示浏览器窗口
    options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36")
    
    driver = webdriver.Chrome(options=options)
    driver.get(url)

    try:
        # 等待广告卡片加载完成
        WebDriverWait(driver, 10).until(
            EC.presence_of_element_located((By.CLASS_NAME, "base-ad-preview-card"))
        )
        
        # 滚动页面加载更多广告(可选,根据需求调整滚动次数)
        for _ in range(3):
            driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
            time.sleep(2)
        
        # 提取页面源码并解析
        soup = BeautifulSoup(driver.page_source, 'html.parser')
        ad_previews = soup.find_all('div', class_='base-ad-preview-card')

        for ad_preview in ad_previews:
            image_tag = ad_preview.select_one('.ad-preview__dynamic-dimensions-image')
            image_link = image_tag.get('src') or image_tag.get('data-delayed-url') if image_tag else None
            title = ad_preview.select_one('.text-md').text.strip() if ad_preview.select_one('.text-md') else None
            content = ad_preview.select_one('.commentary__content').text.strip() if ad_preview.select_one('.commentary__content') else None

            print("Image Link:", image_link)
            print("Title:", title)
            print("Content:", content)
            print("-" * 50)
    finally:
        driver.quit()

LinkedInAdScraperSelenium("SalesForce")

注意事项

  • LinkedIn的页面结构和反爬规则可能随时更新,需定期检查元素类名、属性是否变化。
  • 爬取内容需遵守LinkedIn用户协议,避免高频请求导致IP或账号被封禁。

内容的提问来源于stack exchange,提问作者AaravS

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 06:11:20