LinkedIn Ads Library爬取问题:无法获取图片src属性
解决LinkedIn Ads Library图片及广告内容爬取问题
问题分析
LinkedIn Ads Library的部分内容(包括图片链接)依赖JavaScript动态渲染,直接用requests获取的静态HTML中,图片标签常缺失src属性,真实链接会存储在data-*类属性中;同时,未携带合法请求头的请求易被LinkedIn拦截,导致返回内容不全。
改进现有BeautifulSoup方案
1. 补充请求头模拟浏览器行为
添加真实浏览器的User-Agent,避免被LinkedIn的反爬机制拦截,提升内容获取的完整性。
2. 修正图片属性提取逻辑
实际页面中,广告图片的真实链接通常存储在data-delayed-url属性中,调整提取逻辑优先读取该属性,再降级到src。
改进后的代码:
import requests from bs4 import BeautifulSoup def LinkedInAdScraper(company: str): url = f"https://www.linkedin.com/ad-library/search?accountOwner={company}&dateOption=last-30-days" # 模拟浏览器请求头 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } response = requests.get(url, headers=headers) if response.status_code == 200: soup = BeautifulSoup(response.content, 'html.parser') ad_previews = soup.find_all('div', class_='base-ad-preview-card') for ad_preview in ad_previews: # 提取图片链接:优先读取data-delayed-url,无则 fallback 到src image_tag = ad_preview.select_one('.ad-preview__dynamic-dimensions-image') image_link = image_tag.get('data-delayed-url') or image_tag.get('src') if image_tag else None # 提取标题和内容 title = ad_preview.select_one('.text-md').text.strip() if ad_preview.select_one('.text-md') else None content = ad_preview.select_one('.commentary__content').text.strip() if ad_preview.select_one('.commentary__content') else None print("Image Link:", image_link) print("Title:", title) print("Content:", content) print("-" * 50) else: print(f"请求失败,状态码: {response.status_code}") LinkedInAdScraper("SalesForce")
动态渲染内容爬取方案(Selenium)
若上述方法仍无法获取完整内容(比如部分广告为异步加载),可使用Selenium模拟浏览器渲染页面,确保所有动态内容加载完成后再提取:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time from bs4 import BeautifulSoup def LinkedInAdScraperSelenium(company: str): url = f"https://www.linkedin.com/ad-library/search?accountOwner={company}&dateOption=last-30-days" # 初始化Chrome浏览器(需提前安装ChromeDriver并配置环境变量) options = webdriver.ChromeOptions() options.add_argument("--headless=new") # 无头模式,不显示浏览器窗口 options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36") driver = webdriver.Chrome(options=options) driver.get(url) try: # 等待广告卡片加载完成 WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, "base-ad-preview-card")) ) # 滚动页面加载更多广告(可选,根据需求调整滚动次数) for _ in range(3): driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") time.sleep(2) # 提取页面源码并解析 soup = BeautifulSoup(driver.page_source, 'html.parser') ad_previews = soup.find_all('div', class_='base-ad-preview-card') for ad_preview in ad_previews: image_tag = ad_preview.select_one('.ad-preview__dynamic-dimensions-image') image_link = image_tag.get('src') or image_tag.get('data-delayed-url') if image_tag else None title = ad_preview.select_one('.text-md').text.strip() if ad_preview.select_one('.text-md') else None content = ad_preview.select_one('.commentary__content').text.strip() if ad_preview.select_one('.commentary__content') else None print("Image Link:", image_link) print("Title:", title) print("Content:", content) print("-" * 50) finally: driver.quit() LinkedInAdScraperSelenium("SalesForce")
注意事项
- LinkedIn的页面结构和反爬规则可能随时更新,需定期检查元素类名、属性是否变化。
- 爬取内容需遵守LinkedIn用户协议,避免高频请求导致IP或账号被封禁。
内容的提问来源于stack exchange,提问作者AaravS
相关产品推荐
相关产品推荐

