为何aiohttp.ClientSession.get与urllib.request.urlopen返回内容不同?
问题
我写了一个从URL抓取HTML页面的简单函数:用urllib.request.urlopen请求指定GitHub页面时能正常返回HTML源码,但用aiohttp.ClientSession.get请求同一URL时,返回的却是包含payload、title、locale键的JSON对象。请问我的aiohttp使用哪里出问题了?
urllib代码示例
import urllib.request def download(url): with urllib.request.urlopen(url) as resp: return resp.read(100) url = 'https://github.com/python/cpython/blob/main/Lib/urllib/request.py' print(download(url))
返回结果
> python test_urllib.py b'\n\n\n\n\n\n<!DOCTYPE html>\n<html lang="en" data-color-mode="auto" data-light-theme="light" data-dark-them'
aiohttp代码示例
import aiohttp import asyncio async def download(url): async with aiohttp.ClientSession() as session: async with session.get(url) as response: return await response.content.read(100) url = 'https://github.com/python/cpython/blob/main/Lib/urllib/request.py' print(asyncio.run(download(url)))
返回结果
> python test_aiohttp.py b'{"payload":{"allShortcutsEnabled":false,"fileTree":{"Lib/urllib":{"items":[{"name":"__init__.py","pa'
解决方案
问题出在请求头的User-Agent字段:
urllib.request.urlopen默认会带上类似Python-urllib/3.x的User-Agent,GitHub服务器会将其识别为合法请求,返回完整HTML页面。aiohttp默认的User-Agent是Python/3.x aiohttp/3.x.x,被GitHub反爬机制判定为非浏览器请求,因此返回供前端渲染用的JSON数据(而非直接返回HTML)。
修改方法很简单,给aiohttp的请求添加模拟浏览器的User-Agent头部:
import aiohttp import asyncio async def download(url): # 模拟Chrome浏览器的User-Agent headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } async with aiohttp.ClientSession() as session: async with session.get(url, headers=headers) as response: return await response.content.read(100) url = 'https://github.com/python/cpython/blob/main/Lib/urllib/request.py' print(asyncio.run(download(url)))
修改后运行,就能得到和urllib一致的HTML开头内容。
内容的提问来源于stack exchange,提问作者user2283347
相关产品推荐
相关产品推荐

