You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

新手求助:如何爬取Stack Overflow用户主页的user-tags标签

问题分析与解决

你当前的代码爬取的是Stack Overflow的用户列表页,这个页面本身并不包含用户的user-tags——这类标签仅存在于每个用户的个人主页中。另外直接用requests裸请求容易被反爬拦截,需要模拟浏览器请求头,还得分两步完成爬取:先从列表页抓取用户个人主页的链接,再逐个进入主页提取标签。

修正后的代码

import requests
from bs4 import BeautifulSoup
import time

# 模拟浏览器请求头,规避基础反爬
HEADERS = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
}

# 从用户列表页提取所有用户的个人主页链接
def get_user_profile_urls(list_page_url):
    response = requests.get(list_page_url, headers=HEADERS)
    soup = BeautifulSoup(response.text, 'lxml')
    user_links = []
    # 定位列表中用户的主页链接元素
    user_items = soup.find_all('a', class_='s-user-card--link')
    for item in user_items:
        profile_url = f"https://stackoverflow.com{item['href']}"
        user_links.append(profile_url)
    return user_links

# 从用户个人主页提取user-tags
def scrape_user_tags_from_profile(profile_url):
    try:
        # 添加请求延迟,避免触发反爬机制
        time.sleep(1)
        response = requests.get(profile_url, headers=HEADERS)
        soup = BeautifulSoup(response.text, 'lxml')
        user_tags = []
        # 定位user-tags的容器
        user_tags_div = soup.find('div', class_='user-tags')
        if user_tags_div:
            tags = user_tags_div.find_all('a', class_='post-tag')
            user_tags = [tag.text.strip() for tag in tags]
        return profile_url, user_tags
    except Exception as e:
        print(f"爬取{profile_url}出错: {str(e)}")
        return profile_url, []

# 主执行逻辑
if __name__ == '__main__':
    list_page_url = 'https://stackoverflow.com/users?tab=Reputation&filter=all'
    print("开始爬取用户列表...")
    user_profile_urls = get_user_profile_urls(list_page_url)
    print(f"找到{len(user_profile_urls)}个用户主页链接")
    
    print("\n开始爬取每个用户的标签...")
    for profile_url in user_profile_urls:
        url, tags = scrape_user_tags_from_profile(profile_url)
        if tags:
            print(f"用户主页: {url}\n标签: {', '.join(tags)}\n")
        else:
            print(f"用户主页: {url} 未找到标签\n")
    
    print("爬取完成")

关键注意点

  • 不要缩短请求延迟,time.sleep(1)是最低限度,爬取量较大时建议延长至2-3秒
  • 如果仍被拦截,可以补充Accept-Language等更多请求头字段,或使用代理IP
  • Stack Overflow的页面类名可能会更新,若后续爬取失效,需检查页面元素的class属性是否变化

内容的提问来源于stack exchange,提问作者Bbs

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 05:37:08