如何用BeautifulSoup或Regex提取href中page参数的数值?
提取Goodreads下一页链接中的页码:两种可行方案
当然可以!不管是用BeautifulSoup配合URL解析工具,还是直接用正则表达式,都能轻松提取出href中page=后的数字。下面是具体的实现方法:
方法一:BeautifulSoup + 标准库URL解析(更健壮)
这种方法依赖Python标准库的urllib.parse模块来解析URL参数,不用自己写正则,对URL结构的变化适应性更强——哪怕参数顺序调整、新增其他参数,都不会影响提取结果:
from bs4 import BeautifulSoup import requests from urllib.parse import urlparse, parse_qs request = requests.get('https://www.goodreads.com/quotes/tag/fun?page=1') soup = BeautifulSoup(request.text, 'html.parser') findNext = soup.find("a", class_="next_page") if findNext: # 获取a标签的href属性值 href = findNext.get('href') # 解析URL的查询参数部分 parsed_url = urlparse(href) # 提取page参数的值(返回的是列表,取第一个元素) page_num = parse_qs(parsed_url.query).get('page', [''])[0] if page_num: print(f"提取到的页码:{page_num}") else: print("未找到page参数") else: print("没有找到下一页链接")
方法二:正则表达式(简洁直接)
如果确定Goodreads的URL格式比较固定(比如page=后面一定是数字,且不会有其他类似的参数干扰),用正则表达式会更简洁直观:
from bs4 import BeautifulSoup import requests import re request = requests.get('https://www.goodreads.com/quotes/tag/fun?page=1') soup = BeautifulSoup(request.text, 'html.parser') findNext = soup.find("a", class_="next_page") if findNext: href = findNext.get('href') # 匹配page=后面的连续数字 match = re.search(r'page=(\d+)', href) if match: page_num = match.group(1) print(f"提取到的页码:{page_num}") else: print("未找到page参数") else: print("没有找到下一页链接")
两种方法都能处理你给出的示例输出<a class="next_page" href="/quotes/tag/fun?page=2" rel="next">next »</a>,成功提取出数字2。你可以根据自己的需求选择合适的方案~
内容的提问来源于stack exchange,提问作者Bob Hopez
相关产品推荐
相关产品推荐

