You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何识别状态码200但返回无效内容的网页?

针对你遇到的「状态码200但内容无效」的爬取问题,可以从以下几个维度入手识别这类URL:

1. 对比最终URL与原始请求URL

很多无效请求会被悄悄重定向到404/错误页面,虽然状态码返回200,但最终URL会暴露问题。比如你举的例子里,最终URL包含404-page-not-found或page-not-found这类特征词。

实现思路:

  • 发起请求后,对比response.url和原始请求的URL
  • 检查最终URL是否包含常见错误页关键词,同时判断请求文件类型(如PDF)与最终URL后缀是否匹配

代码示例:

import requests

def is_redirected_to_error(original_url, response):
    error_keywords = {'404', 'page-not-found', 'pagenotfound', 'error'}
    final_url = response.url.lower()
    # 检查最终URL是否含错误关键词
    if any(keyword in final_url for keyword in error_keywords):
        return True
    # 额外判断:如果请求的是PDF,最终URL不是PDF后缀(排除合法重定向情况)
    if original_url.lower().endswith('.pdf') and not final_url.lower().endswith('.pdf'):
        return True
    return False

# 使用示例
original_url = 'http://www.canfor.com/_resources/company/code_of_conduct_2006.pdf'
r = requests.get(original_url, allow_redirects=True)
if is_redirected_to_error(original_url, r):
    print(f"无效URL:{original_url}(重定向到错误页)")

2. 验证响应内容类型是否匹配预期

你请求的是PDF,但返回的是text/html,这种类型不匹配的情况直接说明内容无效。

实现思路:

  • 根据原始URL的后缀(如.pdf、.docx)确定预期的Content-Type
  • 对比响应头中的Content-Type(忽略编码参数,比如application/pdf; charset=utf-8仍属于PDF类型)

代码示例:

def check_content_type(original_url, response):
    # 映射文件后缀到预期Content-Type
    expected_types = {
        '.pdf': 'application/pdf',
        '.docx': 'application/vnd.openxmlformats-officedocument.wordprocessingml.document',
        '.txt': 'text/plain'
    }
    # 获取原始URL的后缀并匹配预期类型
    for ext, content_type in expected_types.items():
        if original_url.lower().endswith(ext):
            actual_type = response.headers.get('content-type', '').split(';')[0].strip()
            if actual_type != content_type:
                return False
    return True

# 使用示例
if not check_content_type(original_url, r):
    print(f"无效URL:{original_url}(内容类型不匹配)")

3. 检测响应内容中的无效关键词

像返回“We'll be back shortly”、“technical issues”这类提示页面,直接在响应文本中搜索这些关键词即可识别。

实现思路:

  • 整理常见的无效内容关键词列表
  • 读取响应文本(优先用响应自带的编码解码)
  • 检查文本中是否包含黑名单关键词

代码示例:

def has_invalid_content(response):
    invalid_keywords = {
        "we'll be back shortly",
        "technical issues",
        "page not found",
        "404 not found",
        "service unavailable"
    }
    try:
        content_text = response.text.lower()
        return any(keyword in content_text for keyword in invalid_keywords)
    except:
        # 解码失败(比如二进制文件),跳过文本检查
        return False

# 使用示例
if has_invalid_content(r):
    print(f"无效URL:{original_url}(内容含无效提示)")

4. 检查响应内容长度的异常值

无效页面(比如404、维护页)的内容长度通常远小于正常文件。比如正常PDF至少几KB,而错误页可能只有几百字节。

实现思路:

  • 根据请求的文件类型设置长度阈值
  • 对比len(response.content)和阈值

代码示例:

def is_content_length_abnormal(original_url, response):
    # 不同类型的阈值设置
    length_thresholds = {
        '.pdf': 1024,  # 1KB
        '.docx': 2048, # 2KB
        '.txt': 512    # 512字节
    }
    for ext, threshold in length_thresholds.items():
        if original_url.lower().endswith(ext):
            if len(response.content) < threshold:
                return True
    return False

# 使用示例
if is_content_length_abnormal(original_url, r):
    print(f"无效URL:{original_url}(内容长度异常)")

5. 整合所有判断逻辑

把上面的方法整合到一个函数里,综合判断URL是否有效:

def is_valid_url(original_url):
    try:
        r = requests.get(original_url, timeout=10)
        if r.status_code != 200:
            return False
        # 依次检查各维度
        if is_redirected_to_error(original_url, r):
            return False
        if not check_content_type(original_url, r):
            return False
        if has_invalid_content(r):
            return False
        if is_content_length_abnormal(original_url, r):
            return False
        return True
    except (requests.exceptions.ConnectionError, requests.exceptions.Timeout):
        return False

# 批量处理URL列表
url_list = [
    'http://www.canfor.com/_resources/company/code_of_conduct_2006.pdf',
    'http://www.coned.com/documents/Con_Edison_2008_Annual_Report.pdf',
    # 其他URL...
]

for url in url_list:
    if is_valid_url(url):
        # 保存有效内容
        filename = url.split('/')[-1]
        with open(f"saved_{filename}", 'wb') as f:
            f.write(requests.get(url).content)
    else:
        print(f"跳过无效URL:{url}")

内容的提问来源于stack exchange,提问作者wavingtide

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 04:01:27