循环中requests.get方法无法正常工作的原因排查求助
循环中requests.get失效的原因排查及修复
以下代码用于批量保存页面HTML,但循环内的requests.get无法正常工作,请排查原因:
import requests from time import sleep import requests from bs4 import BeautifulSoup import csv import time def get_all_pages(): headers = { 'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10.15; rv:109.0) Gecko/20100101 Firefox/110.0', 'Accept': '*/*', 'Accept-Language': 'en-US,en;q=0.5', # 'Accept-Encoding': 'gzip, deflate, br', 'Referer': 'https://ndetal.com/', 'Connection': 'keep-alive', 'Sec-Fetch-Dest': 'empty', 'Sec-Fetch-Mode': 'no-cors', 'Sec-Fetch-Site': 'cross-site', # Requests doesn't support trailers # 'TE': 'trailers', } params = { 'page-url': 'https://ndetal.com/catalog/zapasnye-chasti/noblelift', 'charset': 'utf-8', 'hittoken': '1676964178_c781bda4eee604512a0a33f4b4847066777b3b8c874b5321c4e64c328a41016f', 'browser-info': 'nb:1:cl:836:ar:1:vf:14qzoz81rywx03264pevhj:fu:0:en:utf-8:la:en-US:v:970:cn:1:dp:0:ls:169832776753:hid:386669881:z:240:i:20230221112312:et:1676964192:c:1:rn:670050072:rqn:3:u:1676964177877528558:w:1425x815:s:1440x900x30:sk:2:wv:2:ds:,,,,,,,,,5658,5658,1,:co:0:cpf:1:eu:0:ns:1676964175443:adb:2:pp:3629563401:rqnl:1:st:1676964192', 't': 'gdpr(14)mc(p-1)clc(0-0-0)rqnt(3)aw(1)ecs(0)fid(400)ti(0)', 'force-urlencoded': '1', } with open('link.txt', 'r') as file: src = file.read() soup = BeautifulSoup(src, 'lxml') pages_count = int(soup.find('ul', class_='pagination').find_all('a')[-2].text) for i in range(1, pages_count + 1): url = f'https://ndetal.com/catalog/zapasnye-chasti/noblelift?page={i}' print(url) r = requests.get(url=url, params=params, headers=headers) with open(f"data/page_{i}.html", 'w') as file: file.write(r.text) time.sleep(10) return pages_count + 1 def main(): pages_count = get_all_pages()
失效原因分析
- 硬编码临时参数过期:
params中的hittoken、browser-info是服务器生成的时效性会话凭证,固定使用旧参数会被服务器直接拒绝请求。 - URL参数冲突:构造的URL已包含
?page={i}分页参数,又通过params传递额外参数,导致最终URL参数结构混乱,服务器无法正确解析请求逻辑。正常分页仅需传递page参数即可,其余参数无需全局复用。 - 请求头配置错误:
Sec-Fetch-Site设为cross-site,但请求的是同域名下的页面,应改为same-origin;Referer指向根域名而非分页页面的上级路径,会触发服务器安全校验拦截。 - 缺少异常处理:未捕获请求过程中的HTTP错误、连接超时等异常,无法定位具体失败原因(如403禁止访问、500服务器错误)。
- 文件写入编码问题:直接用
w模式写入HTML内容,未指定编码,可能因编码不一致导致内容乱码或写入失败。 - 冗余模块导入:重复导入
requests,虽不影响运行,但属于无效代码。
修复后的代码示例
import requests from bs4 import BeautifulSoup import time def get_all_pages(): headers = { 'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10.15; rv:109.0) Gecko/20100101 Firefox/110.0', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8', 'Accept-Language': 'en-US,en;q=0.5', 'Referer': 'https://ndetal.com/catalog/zapasnye-chasti/noblelift', 'Connection': 'keep-alive', 'Sec-Fetch-Dest': 'document', 'Sec-Fetch-Mode': 'navigate', 'Sec-Fetch-Site': 'same-origin', 'Upgrade-Insecure-Requests': '1', } # 读取页面获取总页数 with open('link.txt', 'r', encoding='utf-8') as file: src = file.read() soup = BeautifulSoup(src, 'lxml') pages_count = int(soup.find('ul', class_='pagination').find_all('a')[-2].text) for i in range(1, pages_count + 1): url = f'https://ndetal.com/catalog/zapasnye-chasti/noblelift?page={i}' print(f"正在请求:{url}") try: # 仅传递必要的分页参数,移除过期硬编码params r = requests.get(url=url, headers=headers) r.raise_for_status() # 主动捕获HTTP错误 with open(f"data/page_{i}.html", 'w', encoding='utf-8') as file: file.write(r.text) time.sleep(2) # 合理调整休眠时间,避免触发反爬机制 except Exception as e: print(f"请求第{i}页失败:{str(e)}") continue return pages_count + 1 def main(): pages_count = get_all_pages() if __name__ == "__main__": main()
内容的提问来源于stack exchange,提问作者aspitsin
相关产品推荐
相关产品推荐

