如何在Colly中实现递归爬取?分页爬取异常排查
Colly爬虫仅爬取一层的排查与解决
针对你遇到的Colly爬虫无法递归爬取分页的问题,从Colly核心逻辑出发,以下是关键排查点和解决方法:
1. 检查域名过滤规则
如果Collector设置了AllowedDomains,必须确保分页链接的域名(包括子域名)完全在允许列表内。比如分页链接是https://www.example.com/page/2,但你只添加了example.com,Colly会自动过滤该链接。
修正示例:
c := colly.NewCollector( colly.MaxDepth(10), colly.AllowedDomains("example.com", "www.example.com", "anotherexample.com", "www.anotherexample.com"), )
2. 确保分页链接转换为绝对URL
很多网站的分页链接是相对路径(如/page/2),直接调用Visit(link)会导致请求失败,因为Colly无法解析相对路径的上下文。必须用e.Request.AbsoluteURL(link)转换为绝对URL。
修正示例:
c.OnHTML("//a[@class='next-page']", func(e *colly.HTMLElement) { link := e.Attr("href") absLink := e.Request.AbsoluteURL(link) if absLink != "" { e.Request.Visit(absLink) } })
3. 验证分页回调是否触发
添加日志输出,确认你的XPATH选择器是否真的匹配到了分页元素:
c.OnHTML("//a[@class='next-page']", func(e *colly.HTMLElement) { fmt.Println("触发分页回调,找到链接:", e.Attr("href")) // ... 后续访问逻辑 })
如果没有输出,说明XPATH选择器可能存在隐性问题(比如元素是动态加载的,或class名称有空格/变体),可以调整选择器为更宽松的匹配方式,比如//a[contains(text(), '下一页')]或//a[contains(@class, 'next')]。
4. 排查异步模式的等待逻辑
如果开启了c.Async = true,必须在所有Visit调用后添加c.Wait(),否则程序会提前退出,无法完成后续分页请求:
c.Async = true c.Visit("https://example.com/news") c.Visit("https://anotherexample.com/latest") c.Wait() // 必须添加这一行
5. 检查重复请求过滤(可选排查)
Colly默认会过滤重复URL,如果你的分页链接存在参数重复或URL格式问题(比如带无用的session参数),可以临时关闭缓存排查:
c := colly.NewCollector( colly.MaxDepth(10), colly.DisableCache(), // 仅用于排查,生产环境不推荐 )
完整修正示例代码
package main import ( "fmt" "github.com/gocolly/colly/v2" ) func main() { c := colly.NewCollector( colly.MaxDepth(10), colly.AllowedDomains("example.com", "www.example.com", "anotherexample.com", "www.anotherexample.com"), ) // 输出请求日志,方便跟踪深度和URL c.OnRequest(func(r *colly.Request) { fmt.Printf("正在访问:%s(深度:%d)\n", r.URL, r.Depth) }) // 处理文章内容 c.OnHTML("article", func(e *colly.HTMLElement) { title := e.ChildText("h1") fmt.Printf("抓取到文章:%s\n", title) }) // 处理分页链接 c.OnHTML("//a[contains(@class, 'next') or text()='下一页']", func(e *colly.HTMLElement) { link := e.Attr("href") absLink := e.Request.AbsoluteURL(link) if absLink != "" { fmt.Printf("跟进分页链接:%s\n", absLink) if err := e.Request.Visit(absLink); err != nil { fmt.Printf("访问分页链接失败:%v\n", err) } } }) // 输出错误日志 c.OnError(func(r *colly.Response, err error) { fmt.Printf("请求失败:%s,错误信息:%v\n", r.Request.URL, err) }) // 启动爬虫 c.Visit("https://example.com/news") c.Visit("https://anotherexample.com/latest") }
内容的提问来源于stack exchange,提问作者Rahul Seetharaman
相关产品推荐
相关产品推荐

