You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何实现Hacker News多页遍历爬取并保留原有代码结构?

实现思路
  • 完全保留原有的筛选、排序核心逻辑,仅新增分页循环相关代码
  • 把单页的请求、解析操作封装为可复用函数,接收页码作为入参
  • 循环调用单页爬取函数,把多页结果汇总后统一排序输出
  • 可选新增请求延时,避免短时间请求过频被站点限制
修改后完整代码
# 导入所需库
import requests
from bs4 import BeautifulSoup 
import pprint
import time  # 可选,用于控制请求频率

def sort_stories_by_votes(hnlist):  
    # 按点赞数倒序排序帖子
    return sorted(hnlist, key= lambda k:k['votes'], reverse=True)

def create_custom_hn(links, subtext): 
    # 筛选点赞数超过100的帖子
    hn = []
    for idx, item in enumerate(links):
        title = links[idx].getText()
        href = links[idx].get('href', None)
        vote = subtext[idx].select('.score')
        if len(vote):
            points = int(vote[0].getText().replace(' points', ''))
            if points > 99:
                hn.append({'title': title, 'link': href, 'votes': points})
    return hn

# 新增:单页爬取函数
def get_single_page(page_num):
    # 分页链接规则:p参数对应页码
    url = f'https://news.ycombinator.com/news?p={page_num}'
    res = requests.get(url)
    soup = BeautifulSoup(res.text, 'html.parser')
    links = soup.select('.titlelink')
    subtext = soup.select('.subtext')
    return create_custom_hn(links, subtext)

# 新增:批量爬取多页
all_hn = []
# 如需包含当前第一页+接下来10页,用range(1, 12);如需仅爬接下来的10页不含第一页,用range(2,12)
for page in range(1, 12):
    page_data = get_single_page(page)
    all_hn.extend(page_data)
    time.sleep(1)  # 可选:每页间隔1秒请求,避免被限制

# 所有页数据汇总后统一排序输出
pprint.pprint(sort_stories_by_votes(all_hn))
注意事项
  • 请求间隔可根据自身需求调整,删除time.sleep(1)即可无间隔请求,更建议保留避免被站点临时封禁IP
  • 如果站点后续调整了DOM类名,只需要对应修改select里的类名即可,核心逻辑无需变动

内容的提问来源于stack exchange,提问作者TimMTech93

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 19:54:03