You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python+Selenium+BeautifulSoup爬取LinkedIn帖子评论数返回空列表问题

问题根因
  • 类名拼写错误:原代码中填写的class属性存在输入错误,把social-details-social-counts__item误写为social-details-social counts__item,中间的连字符被替换为空格,导致BeautifulSoup无法匹配到对应元素,最终返回空列表。
  • 动态渲染等待缺失:LinkedIn帖子的互动数据为异步加载内容,未等待元素渲染完成就直接拉取页面源码,会出现元素未生成的情况。
  • 无评论场景未做兼容:原逻辑仅在匹配到评论元素时才会写入数值,没有匹配到元素时不会自动补0,不符合需求。
修复后的代码实现
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
import re
from bs4 import BeautifulSoup

postComments = []

# 先等待评论元素加载完成,最长等待10秒
wait = WebDriverWait(browser, 10)
wait.until(EC.presence_of_all_elements_located((By.CLASS_NAME, "social-details-social-counts__comments")))

src = browser.page_source
soup = BeautifulSoup(src, features="lxml")

# 修正class名拼写
bs4TagsComments = soup.find_all("li", attrs = {"class" : "social-details-social-counts__item social-details-social-counts__comments"})

# 处理空匹配的情况,直接补0
if not bs4TagsComments:
    postComments.append(0)
else:
    for tag in bs4TagsComments:
        # 优化取值逻辑,直接提取元素文本即可,无需转字符串匹配正则
        comment_text = tag.get_text(strip=True)
        list_of_matches = re.findall(r'[,0-9]+', comment_text)
        if list_of_matches:
            last_string = list_of_matches.pop()
            without_comma = last_string.replace(',','')
            commentsCount = int(without_comma)
        else:
            commentsCount = 0
        postComments.append(commentsCount)

print(postComments)
额外优化建议
  • 如果需要批量获取多个帖子的评论数,建议遍历每个帖子容器单独提取,避免跨帖子匹配错误
  • 可定期检查LinkedIn的元素类名是否有更新,平台前端迭代时可能会修改类名规则

内容的提问来源于stack exchange,提问作者Zamena Jaffer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 14:06:03