You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python3.8爬取网页仅获“Loading...”,如何获取真实页面内容?

解决Python读取网页真实内容的问题

问题原因

你遇到的是网站的JS反爬验证:页面返回的是一段用于生成验证Cookie的JavaScript代码,浏览器会自动执行这段代码计算出Cookie,然后重新加载页面获取真实内容;而urllib只会静态获取HTML,不会执行JS,所以拿到的只是加载中的页面。

方法一:用Selenium模拟浏览器渲染

Selenium可以模拟真实浏览器的行为,自动执行JS并完成验证,直接获取渲染后的页面内容。

步骤:

  1. 安装Selenium和对应浏览器的驱动(比如ChromeDriver)
  2. 编写代码模拟浏览器访问:
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
import time

# 配置Chrome选项,可选无头模式(不显示浏览器窗口)
chrome_options = Options()
chrome_options.add_argument('--headless=new')
chrome_options.add_argument('--disable-gpu')

# 初始化浏览器驱动
driver = webdriver.Chrome(options=chrome_options)
try:
    link = 'https://opencorporates.com/companies/us_fl/P97000018463'
    driver.get(link)
    # 等待页面加载完成(可根据实际情况调整等待时间)
    time.sleep(3)
    # 获取渲染后的HTML内容
    html = driver.page_source
    print(html)
finally:
    # 关闭浏览器
    driver.quit()

方法二:手动执行页面JS生成Cookie,再用requests请求

页面中的JS代码是用来计算验证Cookie的,我们可以提取这段JS,用execjs执行得到Cookie,然后携带Cookie请求页面。

步骤:

  1. 安装requests和PyExecJS
  2. 编写代码:
import requests
import execjs

# 第一步:获取初始页面的JS代码
link = 'https://opencorporates.com/companies/us_fl/P97000018463'
response = requests.get(link)
html = response.text

# 提取JS中的核心函数并修改,让go函数返回Cookie字符串
js_code = """
function leastFactor(n) {
 if (isNaN(n) || !isFinite(n)) return NaN;
 if (typeof phantom !== 'undefined') return 'phantom';
 if (typeof module !== 'undefined' && module.exports) return 'node';
 if (n==0) return 0;
 if (n%1 || n*n<2) return 1;
 if (n%2==0) return 2;
 if (n%3==0) return 3;
 if (n%5==0) return 5;
 var m=Math.sqrt(n);
 for (var i=7;i<=m;i+=30) {
  if (n%i==0)      return i;
  if (n%(i+4)==0)  return i+4;
  if (n%(i+6)==0)  return i+6;
  if (n%(i+10)==0) return i+10;
  if (n%(i+12)==0) return i+12;
  if (n%(i+16)==0) return i+16;
  if (n%(i+22)==0) return i+22;
  if (n%(i+24)==0) return i+24;
 }
 return n;
}
function go() {
 var p=1654161720790; var s=2297856402; var n;
if ((s >> 15) & 1)	p+=
120411494*	16;/*
p+= */else /* 120886108*
*/p-=	5952163*	16;	if ((s >> 1) & 1)
p+=185622079*
2;/*
p+= */else 
p-=/* 120886108*
*/222069557*	2; if ((s >> 12) & 1)/*
*13;
*/p+=/* 120886108*
*/2437023*	15;/*
p+= */else /*
p+= */p-=
111288784*/*
else p-=
*/13;
if ((s >> 7) & 1)/*
p+= */p+=77307883*/*
*13;
*/8;
else /*
p+= */p-=	86418547* 8;if ((s >> 9) & 1)
p+=/* 120886108*
*/175293250* 10;
else 
p-=57423209*/*
else p-=
*/10; p-=2320115331;
 n=leastFactor(p);
return "KEY="+n+"*"+p/n+":"+s+":3577604866:1;path=/;";
}
"""

# 执行JS获取Cookie
ctx = execjs.compile(js_code)
cookie = ctx.call("go")

# 第二步:携带Cookie请求页面
headers = {
    'Cookie': cookie
}
response = requests.get(link, headers=headers)
print(response.text)

注意事项:

  • 方法二中的p和s值是当前请求的专属参数,后续再次请求时需要重新从初始页面提取最新的参数值。
  • 网站可能随时更新反爬逻辑,以上方法需根据实际情况调整。
  • 爬取网站内容前,请遵守网站的robots.txt协议及相关法律法规。

内容的提问来源于stack exchange,提问作者FrdXt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 07:54:52