You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python网页爬取:世界银行经理页面Div类不一致的解决方法

解决世界银行经理传记爬取的页面结构不一致问题

核心思路

针对不同页面结构不统一的问题,放弃依赖单一固定选择器,改用语义定位+多选择器 fallback+内容过滤的组合策略,确保适配不同页面的传记内容提取。

具体解决方法

  • 语义定位传记区域:先找到页面中标记"Biography"(传记)的标题元素(通常是h2/h3标签),再提取该标题之后的相邻内容块——页面的语义标识一般不会随意变动,这是最可靠的方式。
  • 多选择器降级尝试:预先定义多个可能的传记容器选择器,依次尝试匹配,直到获取到有效内容,覆盖不同页面的结构差异。
  • 过滤无关文本:提取段落文本后,通过文本长度、关键词过滤掉页脚、导航等无关内容,避免混入无效数据。

修改后的代码示例

import requests
import bs4

def get_person_profile(url):
    person_content = requests.get(url).content
    person_soup = bs4.BeautifulSoup(person_content,'html.parser')
    
    # 策略1: 基于语义定位传记区域
    biography_title = person_soup.find(['h2', 'h3'], text=lambda t: t and 'Biography' in t.strip())
    if biography_title:
        profile_text = []
        sibling = biography_title.next_sibling
        while sibling:
            # 遇到下一级标题就停止遍历
            if sibling.name in ['h2', 'h3']:
                break
            # 提取有效长度的段落文本
            if sibling.name == 'p' and len(sibling.text.strip()) > 20:
                profile_text.append(sibling.text.strip())
            # 处理嵌套在div里的段落
            if sibling.name == 'div':
                for p in sibling.find_all('p', text=lambda t: t and len(t.strip())>20):
                    profile_text.append(p.text.strip())
            sibling = sibling.next_sibling
        if profile_text:
            return '\n'.join(profile_text)
    
    # 策略2: 多选择器fallback尝试
    possible_selectors = [
        'div.redesign_text_collapse_expand.parbase.section',
        'div.parbase.section',
        'div.text-long.formatted'
    ]
    for selector in possible_selectors:
        div_data = person_soup.select_one(selector)
        if div_data:
            # 过滤短文本和无关内容
            profile_data = div_data.find_all('p', text=lambda t: t and len(t.strip())>20)
            profile_text = [p.text.strip() for p in profile_data]
            if profile_text:
                return '\n'.join(profile_text)
    
    # 兜底:如果以上都失败,返回提示
    return "无法提取传记内容"

代码说明

  1. 语义定位部分优先查找带有"Biography"的标题,然后遍历后续兄弟元素,提取有效段落(过滤短文本),同时处理嵌套在div里的段落。
  2. 多选择器fallback定义了几个在目标页面中可能出现的容器选择器,依次尝试匹配,避免单一选择器失效。
  3. 所有提取步骤都加入了文本长度过滤(>20字符),排除导航、版权等无关的短文本。

内容的提问来源于stack exchange,提问作者Jairam Jidgekar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 21:11:39