如何识别状态码200但返回无效内容的网页?
针对你遇到的「状态码200但内容无效」的爬取问题,可以从以下几个维度入手识别这类URL:
1. 对比最终URL与原始请求URL
很多无效请求会被悄悄重定向到404/错误页面,虽然状态码返回200,但最终URL会暴露问题。比如你举的例子里,最终URL包含404-page-not-found或page-not-found这类特征词。
实现思路:
- 发起请求后,对比
response.url和原始请求的URL - 检查最终URL是否包含常见错误页关键词,同时判断请求文件类型(如PDF)与最终URL后缀是否匹配
代码示例:
import requests def is_redirected_to_error(original_url, response): error_keywords = {'404', 'page-not-found', 'pagenotfound', 'error'} final_url = response.url.lower() # 检查最终URL是否含错误关键词 if any(keyword in final_url for keyword in error_keywords): return True # 额外判断:如果请求的是PDF,最终URL不是PDF后缀(排除合法重定向情况) if original_url.lower().endswith('.pdf') and not final_url.lower().endswith('.pdf'): return True return False # 使用示例 original_url = 'http://www.canfor.com/_resources/company/code_of_conduct_2006.pdf' r = requests.get(original_url, allow_redirects=True) if is_redirected_to_error(original_url, r): print(f"无效URL:{original_url}(重定向到错误页)")
2. 验证响应内容类型是否匹配预期
你请求的是PDF,但返回的是text/html,这种类型不匹配的情况直接说明内容无效。
实现思路:
- 根据原始URL的后缀(如
.pdf、.docx)确定预期的Content-Type - 对比响应头中的
Content-Type(忽略编码参数,比如application/pdf; charset=utf-8仍属于PDF类型)
代码示例:
def check_content_type(original_url, response): # 映射文件后缀到预期Content-Type expected_types = { '.pdf': 'application/pdf', '.docx': 'application/vnd.openxmlformats-officedocument.wordprocessingml.document', '.txt': 'text/plain' } # 获取原始URL的后缀并匹配预期类型 for ext, content_type in expected_types.items(): if original_url.lower().endswith(ext): actual_type = response.headers.get('content-type', '').split(';')[0].strip() if actual_type != content_type: return False return True # 使用示例 if not check_content_type(original_url, r): print(f"无效URL:{original_url}(内容类型不匹配)")
3. 检测响应内容中的无效关键词
像返回“We'll be back shortly”、“technical issues”这类提示页面,直接在响应文本中搜索这些关键词即可识别。
实现思路:
- 整理常见的无效内容关键词列表
- 读取响应文本(优先用响应自带的编码解码)
- 检查文本中是否包含黑名单关键词
代码示例:
def has_invalid_content(response): invalid_keywords = { "we'll be back shortly", "technical issues", "page not found", "404 not found", "service unavailable" } try: content_text = response.text.lower() return any(keyword in content_text for keyword in invalid_keywords) except: # 解码失败(比如二进制文件),跳过文本检查 return False # 使用示例 if has_invalid_content(r): print(f"无效URL:{original_url}(内容含无效提示)")
4. 检查响应内容长度的异常值
无效页面(比如404、维护页)的内容长度通常远小于正常文件。比如正常PDF至少几KB,而错误页可能只有几百字节。
实现思路:
- 根据请求的文件类型设置长度阈值
- 对比
len(response.content)和阈值
代码示例:
def is_content_length_abnormal(original_url, response): # 不同类型的阈值设置 length_thresholds = { '.pdf': 1024, # 1KB '.docx': 2048, # 2KB '.txt': 512 # 512字节 } for ext, threshold in length_thresholds.items(): if original_url.lower().endswith(ext): if len(response.content) < threshold: return True return False # 使用示例 if is_content_length_abnormal(original_url, r): print(f"无效URL:{original_url}(内容长度异常)")
5. 整合所有判断逻辑
把上面的方法整合到一个函数里,综合判断URL是否有效:
def is_valid_url(original_url): try: r = requests.get(original_url, timeout=10) if r.status_code != 200: return False # 依次检查各维度 if is_redirected_to_error(original_url, r): return False if not check_content_type(original_url, r): return False if has_invalid_content(r): return False if is_content_length_abnormal(original_url, r): return False return True except (requests.exceptions.ConnectionError, requests.exceptions.Timeout): return False # 批量处理URL列表 url_list = [ 'http://www.canfor.com/_resources/company/code_of_conduct_2006.pdf', 'http://www.coned.com/documents/Con_Edison_2008_Annual_Report.pdf', # 其他URL... ] for url in url_list: if is_valid_url(url): # 保存有效内容 filename = url.split('/')[-1] with open(f"saved_{filename}", 'wb') as f: f.write(requests.get(url).content) else: print(f"跳过无效URL:{url}")
内容的提问来源于stack exchange,提问作者wavingtide
相关产品推荐
相关产品推荐

