You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

能否通过Google API下载「已抓取 - 目前未编入索引」的URL列表?

可以通过Google API下载「已抓取 - 目前未编入索引」的URL列表

完全可以,不用依赖界面手动下载,通过Google Search Console API就能批量获取这类URL,核心是用URL Inspection API或searchconsole.urls.list方法配合筛选条件实现:

  • 第一步:先在Google Cloud平台启用Google Search Console API,完成身份验证(支持OAuth2授权或服务账号两种方式,按需选择)。
  • 第二步:调用API时指定目标站点的资源路径(比如sc-domain:yourdomain.com或https://www.yourdomain.com/),设置筛选参数匹配coverageState为"CRAWLED_NOT_INDEXED"的URL——这个状态正好对应界面里「已抓取 - 目前未编入索引」的分类。
  • 第三步:API会返回符合条件的URL列表,可直接批量导出或处理,无需手动操作界面。

举个简单的Python代码片段(基于官方google-api-python-client库):

from googleapiclient.discovery import build
from google.oauth2.credentials import Credentials

# 加载已完成验证的凭据
creds = Credentials.from_authorized_user_file('token.json')
service = build('searchconsole', 'v1', credentials=creds)

# 指定站点并设置筛选条件
site_url = "sc-domain:yourdomain.com"
request = {
    "siteUrl": site_url,
    "filter": {
        "operator": "equals",
        "property": "inspectionResult.indexStatusResult.coverageState",
        "value": "CRAWLED_NOT_INDEXED"
    },
    "pageSize": 1000  # 单页最多返回1000条,可分页获取更多
}

# 调用API并提取URL列表
response = service.urls().list(body=request).execute()
crawled_not_indexed_urls = [item['url'] for item in response.get('urlRows', [])]

如果URL数量超过单页限制,需要通过nextPageToken参数分页获取全部结果。

内容的提问来源于stack exchange,提问作者conteh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 05:52:33