Python3.8爬取网页仅获“Loading...”,如何获取真实页面内容?
解决Python读取网页真实内容的问题
问题原因
你遇到的是网站的JS反爬验证:页面返回的是一段用于生成验证Cookie的JavaScript代码,浏览器会自动执行这段代码计算出Cookie,然后重新加载页面获取真实内容;而urllib只会静态获取HTML,不会执行JS,所以拿到的只是加载中的页面。
方法一:用Selenium模拟浏览器渲染
Selenium可以模拟真实浏览器的行为,自动执行JS并完成验证,直接获取渲染后的页面内容。
步骤:
- 安装Selenium和对应浏览器的驱动(比如ChromeDriver)
- 编写代码模拟浏览器访问:
from selenium import webdriver from selenium.webdriver.chrome.options import Options import time # 配置Chrome选项,可选无头模式(不显示浏览器窗口) chrome_options = Options() chrome_options.add_argument('--headless=new') chrome_options.add_argument('--disable-gpu') # 初始化浏览器驱动 driver = webdriver.Chrome(options=chrome_options) try: link = 'https://opencorporates.com/companies/us_fl/P97000018463' driver.get(link) # 等待页面加载完成(可根据实际情况调整等待时间) time.sleep(3) # 获取渲染后的HTML内容 html = driver.page_source print(html) finally: # 关闭浏览器 driver.quit()
方法二:手动执行页面JS生成Cookie,再用requests请求
页面中的JS代码是用来计算验证Cookie的,我们可以提取这段JS,用execjs执行得到Cookie,然后携带Cookie请求页面。
步骤:
- 安装
requests和PyExecJS - 编写代码:
import requests import execjs # 第一步:获取初始页面的JS代码 link = 'https://opencorporates.com/companies/us_fl/P97000018463' response = requests.get(link) html = response.text # 提取JS中的核心函数并修改,让go函数返回Cookie字符串 js_code = """ function leastFactor(n) { if (isNaN(n) || !isFinite(n)) return NaN; if (typeof phantom !== 'undefined') return 'phantom'; if (typeof module !== 'undefined' && module.exports) return 'node'; if (n==0) return 0; if (n%1 || n*n<2) return 1; if (n%2==0) return 2; if (n%3==0) return 3; if (n%5==0) return 5; var m=Math.sqrt(n); for (var i=7;i<=m;i+=30) { if (n%i==0) return i; if (n%(i+4)==0) return i+4; if (n%(i+6)==0) return i+6; if (n%(i+10)==0) return i+10; if (n%(i+12)==0) return i+12; if (n%(i+16)==0) return i+16; if (n%(i+22)==0) return i+22; if (n%(i+24)==0) return i+24; } return n; } function go() { var p=1654161720790; var s=2297856402; var n; if ((s >> 15) & 1) p+= 120411494* 16;/* p+= */else /* 120886108* */p-= 5952163* 16; if ((s >> 1) & 1) p+=185622079* 2;/* p+= */else p-=/* 120886108* */222069557* 2; if ((s >> 12) & 1)/* *13; */p+=/* 120886108* */2437023* 15;/* p+= */else /* p+= */p-= 111288784*/* else p-= */13; if ((s >> 7) & 1)/* p+= */p+=77307883*/* *13; */8; else /* p+= */p-= 86418547* 8;if ((s >> 9) & 1) p+=/* 120886108* */175293250* 10; else p-=57423209*/* else p-= */10; p-=2320115331; n=leastFactor(p); return "KEY="+n+"*"+p/n+":"+s+":3577604866:1;path=/;"; } """ # 执行JS获取Cookie ctx = execjs.compile(js_code) cookie = ctx.call("go") # 第二步:携带Cookie请求页面 headers = { 'Cookie': cookie } response = requests.get(link, headers=headers) print(response.text)
注意事项:
- 方法二中的
p和s值是当前请求的专属参数,后续再次请求时需要重新从初始页面提取最新的参数值。 - 网站可能随时更新反爬逻辑,以上方法需根据实际情况调整。
- 爬取网站内容前,请遵守网站的
robots.txt协议及相关法律法规。
内容的提问来源于stack exchange,提问作者FrdXt
相关产品推荐
相关产品推荐

