You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用BeautifulSoup提取维基百科See Also章节的链接?

解决方案

你的代码问题在于目标链接并不在See Also对应的h2标签内部,而是存放在h2标签之后的列表容器div中,你需要定位到h2后再查找它的后序同级节点。
以下是可直接运行的修正代码:

import requests
import re
from bs4 import BeautifulSoup

url_req = "https://en.wikipedia.org/wiki/Privacy_law"
response = requests.get(url=url_req)
soup = BeautifulSoup(response.content, 'html.parser')
see_also_links = []

for headline in soup.find_all('h2'):
    # 匹配h2的文本内容,避免匹配到标签内无关字符
    if re.search(r'see.{0,5}also', headline.text, re.IGNORECASE):
        # 定位h2之后存放See Also链接的div容器(维基统一使用div-col类承载该类列表)
        link_container = headline.find_next('div', class_='div-col')
        for a_tag in link_container.find_all('a', href=True):
            # 补全维基的相对链接为完整可访问URL
            full_url = f"https://en.wikipedia.org{a_tag['href']}"
            see_also_links.append(full_url)
        break

print(see_also_links)

内容的提问来源于stack exchange,提问作者palash

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 06:30:01