使用Python+Selenium+BeautifulSoup爬取LinkedIn帖子评论数返回空列表问题
问题根因
- 类名拼写错误:原代码中填写的class属性存在输入错误,把
social-details-social-counts__item误写为social-details-social counts__item,中间的连字符被替换为空格,导致BeautifulSoup无法匹配到对应元素,最终返回空列表。 - 动态渲染等待缺失:LinkedIn帖子的互动数据为异步加载内容,未等待元素渲染完成就直接拉取页面源码,会出现元素未生成的情况。
- 无评论场景未做兼容:原逻辑仅在匹配到评论元素时才会写入数值,没有匹配到元素时不会自动补0,不符合需求。
修复后的代码实现
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By import re from bs4 import BeautifulSoup postComments = [] # 先等待评论元素加载完成,最长等待10秒 wait = WebDriverWait(browser, 10) wait.until(EC.presence_of_all_elements_located((By.CLASS_NAME, "social-details-social-counts__comments"))) src = browser.page_source soup = BeautifulSoup(src, features="lxml") # 修正class名拼写 bs4TagsComments = soup.find_all("li", attrs = {"class" : "social-details-social-counts__item social-details-social-counts__comments"}) # 处理空匹配的情况,直接补0 if not bs4TagsComments: postComments.append(0) else: for tag in bs4TagsComments: # 优化取值逻辑,直接提取元素文本即可,无需转字符串匹配正则 comment_text = tag.get_text(strip=True) list_of_matches = re.findall(r'[,0-9]+', comment_text) if list_of_matches: last_string = list_of_matches.pop() without_comma = last_string.replace(',','') commentsCount = int(without_comma) else: commentsCount = 0 postComments.append(commentsCount) print(postComments)
额外优化建议
- 如果需要批量获取多个帖子的评论数,建议遍历每个帖子容器单独提取,避免跨帖子匹配错误
- 可定期检查LinkedIn的元素类名是否有更新,平台前端迭代时可能会修改类名规则
内容的提问来源于stack exchange,提问作者Zamena Jaffer
相关产品推荐
相关产品推荐

