You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中如何将BeautifulSoup解析结果转换为可排序的列表?

实现方案

核心逻辑:初始化空列表存储全量数据,每页提取到的数值追加到列表中,全部爬取完成后直接调用列表排序方法即可。

场景1:接口返回HTML页面,需用BeautifulSoup解析

完整代码如下:

from bs4 import BeautifulSoup as bs
import requests

url = "http://challenge.dienekes.com.br/api/numbers?page="
# 初始化空列表存储所有数值
all_numbers = []

for page in range(1, 10000):
    try:
        req = requests.get(url + str(page), timeout=10)
        # 校验请求是否成功
        req.raise_for_status()
        soup = bs(req.text, 'html.parser')
        
        # 请根据实际页面的数值存放标签调整选择器,示例假设数值存放在class为num的p标签中
        num_tags = soup.find_all("p", class_="num")
        for tag in num_tags:
            # 提取文本转整数后加入列表
            num = int(tag.get_text(strip=True))
            all_numbers.append(num)
    except Exception as e:
        print(f"第{page}页爬取失败:{e}")
        continue

# 升序排序,需降序则添加参数reverse=True
all_numbers.sort()
# 输出前10条验证结果
print(all_numbers[:10])

场景2:接口返回JSON格式数值数组(无需使用bs4)

如果接口直接返回JSON格式的数字列表,可直接简化处理,性能更高:

import requests

url = "http://challenge.dienekes.com.br/api/numbers?page="
all_numbers = []

for page in range(1, 10000):
    try:
        resp = requests.get(url + str(page), timeout=10)
        resp.raise_for_status()
        # 直接解析JSON得到当前页数字列表,合并到总列表
        all_numbers.extend(resp.json())
    except Exception as e:
        print(f"第{page}页处理失败:{e}")
        continue

all_numbers.sort()

注意事项

  • 标签选择器需匹配实际页面结构:如果数值存放在其他标签(如li、div)或用id标识,修改find_all的参数即可。如果页面是纯文本每行一个数字,可直接用soup.get_text(strip=True).split()拆分得到字符串列表后转整数。
  • 可按需调整排序逻辑:sort()默认升序,需降序可写为all_numbers.sort(reverse=True)。
  • 建议添加请求头模拟浏览器访问,降低接口拦截概率。

内容的提问来源于stack exchange,提问作者Raoni Duarte

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.23 22:45:08