You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup提取txt格式HTML文件中的人员简介信息

问题根因

你代码中依赖的ng-tns-c82-4511是Angular框架生成的动态随机类后缀,不同人员卡片、不同页面实例的后缀数值都会变化,写死固定值会导致匹配不到元素,是提取失效的核心原因。

调整方案
  • 定位规则只保留固定静态类名,删除所有ng-tns-xxx动态类的匹配逻辑
  • 先遍历所有人员卡片的公共父容器,再在单个父容器内查找对应字段,避免全局查找导致的顺序错乱、匹配错误
  • 提取内容时增加判空逻辑,同时清理多余换行、空格得到干净文本
修正后代码示例
# 先替换下方person-card为你HTML中每个人员外层父容器的公共类名
person_cards = soup.find_all(attrs={"class": "person-card"})

result = []
for card in person_cards:
    # 提取头像链接
    photo_elem = card.find(attrs={"class": "profile-image"})
    photo_src = photo_elem["src"].strip() if photo_elem and photo_elem.has_attr("src") else ""
    
    # 提取姓名
    name_elem = card.find(attrs={"class": "link-is-positive"})
    name = name_elem.get_text(strip=True) if name_elem else ""
    
    # 提取职位
    title_elem = card.find(attrs={"class": "margin--none ng-star-inserted"})
    title = title_elem.get_text(strip=True) if title_elem else ""
    
    # 提取个人介绍
    intro_elem = card.find(attrs={"class": "summary-text margin--none"})
    intro = intro_elem.get_text(strip=True) if intro_elem else ""
    
    result.append({
        "photo_src": photo_src,
        "name": name,
        "title": title,
        "intro": intro
    })
补充说明

如果找不到人员卡片的公共父类,也可以用CSS类名模糊匹配规则调整定位逻辑,比如用card.select_one('[class*="summary-text"]')来匹配所有类名包含summary-text的元素。

内容的提问来源于stack exchange,提问作者Boomshakalaka

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 15:57:04