Java爬虫疑问:如何循环爬取Barnes & Noble网站图书标题
问题分析与解决方案
核心问题:CSS选择器语法错误
你的循环输出为空的根本原因是CSS选择器写法错误,导致无法匹配到目标元素:
1. 列表容器选择器错误
原代码中:
Elements bookList = doc.select("ol.product-shelf-list product-list-view > li");
这里的product-list-view是<ol>标签的第二个class,多个class选择器需要用.连接(空格表示后代元素)。正确写法应该是:
Elements bookList = doc.select("ol.product-shelf-list.product-list-view > li");
2. 标题元素选择器错误
循环内的选择器:
Element titleElement = book.select(".row topX-row > div.product-info-view > div.product-shelf-title > .product-info-view").first();
存在两个错误:
.row topX-row需改为.row.topX-row(两个class用点连接)- 最终要匹配的是标题的
h3.product-info-title,而非.product-info-view
简化后的正确选择器可直接定位标题:
Element titleElement = book.select(".product-info-title").first();
完整修正代码
String url = "https://www.barnesandnoble.com/b/books/_/N-1fZ29Z8q8"; Document doc = null; try { // 添加User-Agent模拟浏览器请求,避免被反爬拦截 doc = Jsoup.connect(url) .userAgent("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36") .get(); } catch ( Exception e) { e.printStackTrace(); } // 输出Top100标题 Element top100 = doc.select(".text").first(); if(top100 != null) { System.out.println(top100.text()); } // 爬取所有图书标题 Elements bookList = doc.select("ol.product-shelf-list.product-list-view > li"); for(Element book : bookList) { Element titleElement = book.select(".product-info-title").first(); if(titleElement != null) { // 防止空指针异常 String title = titleElement.text(); System.out.println(title); } }
额外建议
- 始终添加
User-Agent请求头:多数网站会拦截无标识的爬虫请求,模拟浏览器UA能提升爬取成功率 - 增加空指针判断:避免因页面结构临时变化导致程序崩溃
- 遵守网站
robots.txt规则:爬取前确认网站是否允许爬虫行为,规避法律风险
内容的提问来源于stack exchange,提问作者CHRIS G
相关产品推荐
相关产品推荐

