You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Google Vertex AI搜索高级网站索引中过滤PDF及纯网页结果?

解决Google Vertex AI高级网站索引的文件过滤问题

仅返回PDF文件的方案

由于高级网站索引不支持fileType字段,只能通过官方支持的meta字段实现过滤:

  • 匹配PDF对应的MIME类型:
    {filter: "meta.content_type: \"application/pdf\""}
    
  • 通过文件名后缀过滤:
    {filter: "meta.file_name: \"*.pdf\""}
    

仅返回网页(排除PDF、Word等文件)的方案

同样借助meta字段,指定网页的MIME类型或排除非网页文件类型:

  • 直接匹配HTML网页:
    {filter: "meta.content_type: \"text/html\""}
    
  • 排除多种文件类型(用NOT逻辑组合):
    {filter: "NOT meta.content_type: \"application/pdf\" AND NOT meta.content_type: \"application/msword\" AND NOT meta.content_type: \"application/vnd.openxmlformats-officedocument.wordprocessingml.document\""}
    

注意事项

  • meta字段过滤依赖爬虫抓取到的文件元数据,需确保爬虫配置正确获取Content-Type、文件名等信息
  • 测试过滤规则时,建议先小范围爬取目标站点,验证元数据是否被正确索引

内容的提问来源于stack exchange,提问作者mana

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 03:55:59