You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup或Regex提取href中page参数的数值?

提取Goodreads下一页链接中的页码:两种可行方案

当然可以!不管是用BeautifulSoup配合URL解析工具,还是直接用正则表达式,都能轻松提取出href中page=后的数字。下面是具体的实现方法:

方法一:BeautifulSoup + 标准库URL解析(更健壮)

这种方法依赖Python标准库的urllib.parse模块来解析URL参数,不用自己写正则,对URL结构的变化适应性更强——哪怕参数顺序调整、新增其他参数,都不会影响提取结果:

from bs4 import BeautifulSoup
import requests
from urllib.parse import urlparse, parse_qs

request = requests.get('https://www.goodreads.com/quotes/tag/fun?page=1')
soup = BeautifulSoup(request.text, 'html.parser')
findNext = soup.find("a", class_="next_page")

if findNext:
    # 获取a标签的href属性值
    href = findNext.get('href')
    # 解析URL的查询参数部分
    parsed_url = urlparse(href)
    # 提取page参数的值(返回的是列表,取第一个元素)
    page_num = parse_qs(parsed_url.query).get('page', [''])[0]
    if page_num:
        print(f"提取到的页码:{page_num}")
    else:
        print("未找到page参数")
else:
    print("没有找到下一页链接")

方法二:正则表达式(简洁直接)

如果确定Goodreads的URL格式比较固定(比如page=后面一定是数字,且不会有其他类似的参数干扰),用正则表达式会更简洁直观:

from bs4 import BeautifulSoup
import requests
import re

request = requests.get('https://www.goodreads.com/quotes/tag/fun?page=1')
soup = BeautifulSoup(request.text, 'html.parser')
findNext = soup.find("a", class_="next_page")

if findNext:
    href = findNext.get('href')
    # 匹配page=后面的连续数字
    match = re.search(r'page=(\d+)', href)
    if match:
        page_num = match.group(1)
        print(f"提取到的页码:{page_num}")
    else:
        print("未找到page参数")
else:
    print("没有找到下一页链接")

两种方法都能处理你给出的示例输出<a class="next_page" href="/quotes/tag/fun?page=2" rel="next">next »</a>,成功提取出数字2。你可以根据自己的需求选择合适的方案~

内容的提问来源于stack exchange,提问作者Bob Hopez

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:39:46