You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python requests获取Stack Overflow搜索结果跳转页的HTML内容

代码问题修复与功能实现

原有代码问题

  • 缩进错误:遍历搜索结果的for que in questions循环定义在stackoverflow函数外部,questions属于函数内部局部变量,外部无法访问,运行会直接抛出变量未定义报错
  • 链接提取逻辑错误:不需要通过问题标题反向匹配href属性,.question-hyperlink本身就是带目标链接的a标签,直接提取其href属性即可

修正后可运行代码

import requests
from bs4 import BeautifulSoup

# 配置请求头,避免被反爬拦截
HEADERS = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36"
}

def stackoverflow(question):
    questionAdjusted = question.replace(' ','+')
    # 请求搜索结果页
    search_res = requests.get("https://pt.stackoverflow.com/search?q="+questionAdjusted, headers=HEADERS)
    soup = BeautifulSoup(search_res.text,"html.parser")
    # 直接提取前3条搜索结果
    questions = soup.select(".question-summary")[:3]

    for que in questions:
        link_ele = que.select_one('.question-hyperlink')
        # 拼接完整的问题页面链接
        full_url = "https://pt.stackoverflow.com" + link_ele['href']
        # 请求问题页面获取HTML内容
        page_res = requests.get(full_url, headers=HEADERS)
        page_html = page_res.text  # 该变量即为对应页面的完整HTML内容,可直接用于对接Notion API
        
        # 测试输出,可根据需求删除
        print(f"问题标题:{link_ele.getText()}")
        print(f"问题链接:{full_url}")
        print(f"HTML内容长度:{len(page_html)}\n")

stackoverflow('python database')

功能说明

修正后的代码可直接实现需要的核心逻辑:

  • 传入搜索关键词后自动请求葡萄牙语Stack Overflow搜索结果
  • 自动提取前3条结果的完整访问链接
  • 依次请求每个链接获取完整HTML内容,拿到的page_html变量可直接作为后续对接Notion API的数据源

内容的提问来源于stack exchange,提问作者Kaue Gomes

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 16:27:03