You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用BeautifulSoup爬取LinkedIn公司主页粉丝数

LinkedIn公司页粉丝数爬取方案

我已通过下述代码成功爬取LinkedIn平台的公司名称、所在位置信息,但在爬取公司粉丝数量时遇到障碍。

参考页面HTML结构

<div class="block mt2">
<div>
<h1 class="ember-view t-24 t-black t-bold full-width" id="ember28" title="Pacific Retail Capital Partners">
<span dir="ltr">Pacific Retail Capital Partners</span>
</h1>
<p class="org-top-card-summary__tagline t-16 t-black">
      Our decades of experience and innovative strategies are transforming retail-led centers into high-performing properties.
    </p>
<!-- -->
<div class="org-top-card-summary-info-list t-14 t-black--light">
<div class="org-top-card-summary-info-list__info-item">
      Leasing Non-residential Real Estate
    </div>
<!-- -->
<div class="inline-block">
<div class="org-top-card-summary-info-list__info-item">
        El Segundo, CA
      </div>
<!-- -->
<div class="org-top-card-summary-info-list__info-item">
          4,047 followers
        </div>
</div>
</div>
</div>
</div>

已实现的爬取逻辑

  • 公司名称爬取代码:
info_div = soup.find('div', {'class' : 'block mt2'})
#print(info_div)
info_name = info_div.find_all('h1')
company_name = info_name[0].get_text().strip()
print(company_name, type(company_name),len(company_name))
  • 公司位置爬取代码:
info_block = info_div.find_all('div', {'class' : 'inline-block'})
info_loc = info_block[0].find('div', {'class' : 'org-top-card-summary-info-list__info-item'}).get_text().strip()
print(info_loc)

粉丝数爬取实现

你之前获取位置时,用find()只会匹配容器下第一个符合class的元素,粉丝数是同个inline-block容器下第二个同class的信息项,直接取对应索引即可:

info_block = info_div.find_all('div', {'class' : 'inline-block'})
# 获取容器下所有信息条目
info_items = info_block[0].find_all('div', {'class' : 'org-top-card-summary-info-list__info-item'})
# 索引0是位置,索引1就是粉丝数字段
follower_text = info_items[1].get_text().strip()
print(follower_text)  # 输出:4,047 followers

# 如果需要提取纯数字的粉丝量,做简单字符串处理即可
follower_count = int(follower_text.split()[0].replace(',', ''))
print(follower_count) # 输出:4047

如果要提升代码容错率,避免页面结构微调导致索引错位,可以直接匹配包含followers关键词的元素,不用依赖固定位置:

follower_elem = info_div.find(
    'div',
    class_='org-top-card-summary-info-list__info-item',
    string=lambda t: t and 'followers' in t.lower()
)
follower_text = follower_elem.get_text().strip()

内容的提问来源于stack exchange,提问作者yash agarwal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 00:18:15