You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python for循环分页爬取Lazada网页时URL占位符未替换问题求助

问题原因

  • 你调用url.format(page)后没有将结果赋值回url变量,打印时输出的仍然是未替换占位符的原始模板字符串
  • 你预期输出的数字外层包裹大括号,需要对字符串中的大括号做转义处理(Python 中连续两个大括号{{}}格式化后会输出为单个大括号)
  • 额外优化:HTMLSession不需要在每次循环中重复创建,放在循环外即可减少不必要的资源消耗

修正代码

场景1:实际请求的URL就需要带大括号包裹的页码

from requests_html import HTMLSession
from bs4 import BeautifulSoup

s = HTMLSession()
# 模板中两层大括号为转义,格式化后会保留一层包裹数字
url_template = 'https://www.lazada.com.ph/catalog/?q=laptop&page={{{}}}'

for page in range(1,6):
    final_url = url_template.format(page)
    r = s.get(final_url)
    print(final_url)

场景2:仅打印时需要展示大括号,实际请求用正常分页格式(推荐,符合Lazada实际接口规则)

from requests_html import HTMLSession
from bs4 import BeautifulSoup

s = HTMLSession()
# 实际请求的URL模板不需要转义大括号
url_template = 'https://www.lazada.com.ph/catalog/?q=laptop&page={}'

for page in range(1,6):
    final_url = url_template.format(page)
    r = s.get(final_url)
    # 打印时单独加转义的大括号,不影响实际请求
    print(f"https://www.lazada.com.ph/catalog/?q=laptop&page={{{page}}}")

两种写法运行后都会输出你预期的结果:

https://www.lazada.com.ph/catalog/?q=laptop&page={1}
https://www.lazada.com.ph/catalog/?q=laptop&page={2}
https://www.lazada.com.ph/catalog/?q=laptop&page={3}
https://www.lazada.com.ph/catalog/?q=laptop&page={4}
https://www.lazada.com.ph/catalog/?q=laptop&page={5}

内容的提问来源于stack exchange,提问作者Nate

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 16:51:01