如何用Beautiful Soup获取LinkedIn经验模块工作信息(动态ID解决)
爬取LinkedIn工作经历的问题解决
问题背景
爬取LinkedIn信息时,遇到以下情况:
- 目标section的ID会随页面刷新变化,无法通过固定ID定位
- 使用
soup.find_all(id="experience")只能拿到空的锚点div,无法获取实际工作信息
提供的HTML结构中,id="experience"是一个空的锚点div,实际工作内容在它的兄弟节点内。现有代码:
experiences = soup.find_all("section", {"class": "artdeco-card ember-view relative break-words pb3 mt2"})
解决方案
方法1:通过锚点定位父section,再提取内容
先找到id="experience"的锚点div,再定位到它所在的父section,最后在section内提取工作项:
# 定位experience锚点 experience_anchor = soup.find(id="experience") # 获取锚点所在的父section experience_section = experience_anchor.find_parent("section", class_="artdeco-card ember-view relative break-words pb3 mt2") # 提取所有工作项 job_items = experience_section.find_all("li", class_="artdeco-list__item pvs-list__item--line-separated pvs-list__item--one-column") # 解析每个工作项的信息 for item in job_items: title = item.find("span", class_="mr1 t-bold").get_text(strip=True) company = item.find("span", class_="t-14 t-normal").get_text(strip=True) duration = item.find_all("span", class_="t-14 t-normal t-black--light")[0].get_text(strip=True) location = item.find_all("span", class_="t-14 t-normal t-black--light")[1].get_text(strip=True) print(f"职位: {title}") print(f"公司: {company}") print(f"工作时长: {duration}") print(f"工作地点: {location}") print("-"*30)
方法2:直接定位锚点的兄弟内容容器
因为实际内容在锚点div的下一个兄弟节点(pvs-list__outer-container)中,可以直接定位该容器:
# 定位experience锚点 experience_anchor = soup.find(id="experience") # 获取锚点后的内容容器 content_container = experience_anchor.find_next_sibling("div", class_="pvs-list__outer-container") # 提取工作项 job_items = content_container.find_all("li", class_="artdeco-list__item pvs-list__item--line-separated pvs-list__item--one-column") # 后续解析逻辑同方法1
核心原因
id="experience"的div是LinkedIn页面的锚点标签,本身没有包含工作内容,实际信息存储在它的相邻兄弟节点或父section的子元素中,因此需要通过关联节点定位到内容容器。
内容的提问来源于stack exchange,提问作者Emilio Conde
相关产品推荐
相关产品推荐

