添加User-Agent后爬虫仍遭拦截,返回未授权框架窗口错误求助
Hey there! Let's break down why you're hitting that error even after adding a User-Agent header, and how to fix it.
The Core Issue
You created a Request object with a custom User-Agent, but you never actually used it when making the request! Your code calls uReq(myurl) directly, which uses urllib's default request headers instead of the ones you defined. That's why the site's anti-scraping measures are still blocking you.
Step-by-Step Fixes
- Use your custom Request object: Replace the line
uClient = uReq(myurl)withuClient = uReq(req)so your custom headers are sent with the request. - Add more realistic request headers: To make your request look more like a real browser, add additional headers like
AcceptandRefererto your headers dictionary. This helps bypass basic anti-scraping checks.
Modified Working Code
# Creating my own webcrawler from urllib.request import urlopen as uReq from bs4 import BeautifulSoup as soup import urllib.request myurl = 'https://nvd.nist.gov/vuln/data-feeds' # Define request headers with additional fields for authenticity req = urllib.request.Request( myurl, data=None, headers={ 'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_9_3) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/35.0.1916.47 Safari/537.36', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8', 'Referer': 'https://nvd.nist.gov/' } ) # Opening connection with the custom request (this is the key fix!) uClient = uReq(req) page_html = uClient.read() uClient.close() # HTML parsing page_soup = soup(page_html, 'html.parser') print(page_soup.h1) containers = page_soup.findAll('td', {'rowspan': '1', 'class': 'x-hidden-focus'}) print(len(containers))
Why This Works
The "unauthorized frame window" error is a common anti-scraping check. The site is verifying that the request comes from a legitimate browser context, which requires proper request headers. By using your custom Request object with realistic headers, you're mimicking a real browser visit, so the site allows your request through.
内容的提问来源于stack exchange,提问作者Beni billhardt

