如何用Python requests获取Stack Overflow搜索结果跳转页的HTML内容
代码问题修复与功能实现
原有代码问题
- 缩进错误:遍历搜索结果的
for que in questions循环定义在stackoverflow函数外部,questions属于函数内部局部变量,外部无法访问,运行会直接抛出变量未定义报错 - 链接提取逻辑错误:不需要通过问题标题反向匹配href属性,
.question-hyperlink本身就是带目标链接的a标签,直接提取其href属性即可
修正后可运行代码
import requests from bs4 import BeautifulSoup # 配置请求头,避免被反爬拦截 HEADERS = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" } def stackoverflow(question): questionAdjusted = question.replace(' ','+') # 请求搜索结果页 search_res = requests.get("https://pt.stackoverflow.com/search?q="+questionAdjusted, headers=HEADERS) soup = BeautifulSoup(search_res.text,"html.parser") # 直接提取前3条搜索结果 questions = soup.select(".question-summary")[:3] for que in questions: link_ele = que.select_one('.question-hyperlink') # 拼接完整的问题页面链接 full_url = "https://pt.stackoverflow.com" + link_ele['href'] # 请求问题页面获取HTML内容 page_res = requests.get(full_url, headers=HEADERS) page_html = page_res.text # 该变量即为对应页面的完整HTML内容,可直接用于对接Notion API # 测试输出,可根据需求删除 print(f"问题标题:{link_ele.getText()}") print(f"问题链接:{full_url}") print(f"HTML内容长度:{len(page_html)}\n") stackoverflow('python database')
功能说明
修正后的代码可直接实现需要的核心逻辑:
- 传入搜索关键词后自动请求葡萄牙语Stack Overflow搜索结果
- 自动提取前3条结果的完整访问链接
- 依次请求每个链接获取完整HTML内容,拿到的
page_html变量可直接作为后续对接Notion API的数据源
内容的提问来源于stack exchange,提问作者Kaue Gomes
相关产品推荐
相关产品推荐

