使用Colly爬取赛事数据出错,请求排查Go代码问题
问题排查与修复
你的代码核心问题在于选择器层级错误,以及没有考虑网站类名的动态随机后缀,导致无法定位到正确的赛事元素。以下是具体修复步骤和代码:
问题分析
- 你监听的
div[class='competition_competition_wbjsu']是联赛容器,而非单个赛事的容器,直接在这个层级找球队、比分等字段,会定位不到正确元素。 - 网站的类名带有随机后缀(比如
competition_competition_wbjsu中的wbjsu),固定完整类名的选择器会失效,应该用前缀/包含匹配。 - 没有遍历联赛容器下的单个赛事条目,导致无法获取每个赛事的独立数据。
修复后的代码
// scraper.go package main import ( "fmt" "github.com/gocolly/colly" ) type Fixture struct { League string Winner string Loser string WinScore string LosScore string } func main() { fmt.Println("Starting program") c := colly.NewCollector( colly.AllowedDomains("www.goal.com", "goal.com"), colly.Async(true), // 添加浏览器UA,避免被反爬拦截 colly.UserAgent("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"), ) // 监听联赛容器 c.OnHTML("div[class^='competition_competition']", func(e *colly.HTMLElement) { // 获取联赛名称(用前缀匹配类名,兼容随机后缀) leagueName := e.ChildText("div[class^='competition_name']") // 遍历联赛下的每个赛事条目 e.ForEach("div[class^='match_row']", func(_ int, matchEl *colly.HTMLElement) { fixture := Fixture{ League: leagueName, // 定位获胜球队(包含winner关键字的容器下的h4) Winner: matchEl.ChildText("div[class*='winner'] h4[class^='name_name']"), // 定位失利球队(排除包含winner的容器) Loser: matchEl.ChildText("div[class^='match_team']:not([class*='winner']) h4[class^='name_name']"), // 定位两队比分 WinScore: matchEl.ChildText("div[class^='match_score'] div[class*='team-a'] p"), LosScore: matchEl.ChildText("div[class^='match_score'] div[class*='team-b'] p"), } fmt.Printf("%+v\n", fixture) }) }) c.OnError(func(r *colly.Response, err error) { fmt.Println("Request URL:", r.Request.URL, "failed with response:", r.StatusCode, "\nError:", err) }) url := "https://www.goal.com/en-us/live-scores" c.Visit(url) c.Wait() fmt.Println("End of Program") }
关键修改点
- 选择器优化:用
[class^='xxx'](前缀匹配)和[class*='xxx'](包含匹配)替代固定完整类名,避免因网站类名动态变化失效。 - 层级调整:先定位联赛容器,再遍历其中的每个赛事条目,确保每个赛事数据独立获取。
- 反爬处理:添加浏览器User-Agent,降低被网站拦截的概率。
- 错误信息优化:在OnError中输出响应状态码,便于排查请求问题。
内容的提问来源于stack exchange,提问作者luis
相关产品推荐
相关产品推荐

