Python网页爬取:世界银行经理页面Div类不一致的解决方法
解决世界银行经理传记爬取的页面结构不一致问题
核心思路
针对不同页面结构不统一的问题,放弃依赖单一固定选择器,改用语义定位+多选择器 fallback+内容过滤的组合策略,确保适配不同页面的传记内容提取。
具体解决方法
- 语义定位传记区域:先找到页面中标记"Biography"(传记)的标题元素(通常是h2/h3标签),再提取该标题之后的相邻内容块——页面的语义标识一般不会随意变动,这是最可靠的方式。
- 多选择器降级尝试:预先定义多个可能的传记容器选择器,依次尝试匹配,直到获取到有效内容,覆盖不同页面的结构差异。
- 过滤无关文本:提取段落文本后,通过文本长度、关键词过滤掉页脚、导航等无关内容,避免混入无效数据。
修改后的代码示例
import requests import bs4 def get_person_profile(url): person_content = requests.get(url).content person_soup = bs4.BeautifulSoup(person_content,'html.parser') # 策略1: 基于语义定位传记区域 biography_title = person_soup.find(['h2', 'h3'], text=lambda t: t and 'Biography' in t.strip()) if biography_title: profile_text = [] sibling = biography_title.next_sibling while sibling: # 遇到下一级标题就停止遍历 if sibling.name in ['h2', 'h3']: break # 提取有效长度的段落文本 if sibling.name == 'p' and len(sibling.text.strip()) > 20: profile_text.append(sibling.text.strip()) # 处理嵌套在div里的段落 if sibling.name == 'div': for p in sibling.find_all('p', text=lambda t: t and len(t.strip())>20): profile_text.append(p.text.strip()) sibling = sibling.next_sibling if profile_text: return '\n'.join(profile_text) # 策略2: 多选择器fallback尝试 possible_selectors = [ 'div.redesign_text_collapse_expand.parbase.section', 'div.parbase.section', 'div.text-long.formatted' ] for selector in possible_selectors: div_data = person_soup.select_one(selector) if div_data: # 过滤短文本和无关内容 profile_data = div_data.find_all('p', text=lambda t: t and len(t.strip())>20) profile_text = [p.text.strip() for p in profile_data] if profile_text: return '\n'.join(profile_text) # 兜底:如果以上都失败,返回提示 return "无法提取传记内容"
代码说明
- 语义定位部分优先查找带有"Biography"的标题,然后遍历后续兄弟元素,提取有效段落(过滤短文本),同时处理嵌套在div里的段落。
- 多选择器fallback定义了几个在目标页面中可能出现的容器选择器,依次尝试匹配,避免单一选择器失效。
- 所有提取步骤都加入了文本长度过滤(>20字符),排除导航、版权等无关的短文本。
内容的提问来源于stack exchange,提问作者Jairam Jidgekar
相关产品推荐
相关产品推荐

