Java Selenium爬取Kream网站中途报Connection reset错误求助
Selenium爬取Kream中途触发Connection reset报错修复
问题现象
使用Java+Selenium开发爬虫爬取Kream商品列表,程序启动后初始运行正常,执行到无限滚动爬取流程中途意外终止,已确认Chrome与ChromeDriver均为103版本,版本匹配排除兼容问题。
原实现代码
import java.util.Date; import java.util.List; import org.openqa.selenium.By; import org.openqa.selenium.JavascriptExecutor; import org.openqa.selenium.PageLoadStrategy; import org.openqa.selenium.WebDriver; import org.openqa.selenium.WebElement; import org.openqa.selenium.chrome.ChromeDriver; import org.openqa.selenium.chrome.ChromeOptions; public class Kream { private static WebDriver driver; public static final String WEB_DRIVER_ID = "webdriver.chrome.driver"; public static final String WEB_DRIVER_PATH = "C://chromedriver.exe"; public static void main(String[] args) throws InterruptedException { Kream krm = new Kream(); int rank = 0; System.setProperty(WEB_DRIVER_ID, WEB_DRIVER_PATH); ChromeOptions options = new ChromeOptions(); options.addArguments("headless"); options.setPageLoadStrategy(PageLoadStrategy.NORMAL); WebDriver driver = new ChromeDriver(options); try { String url = "https://kream.co.kr/search?category_id=34&sort=popular&per_page=40"; driver.get(url); List<WebElement> el = driver.findElements(By.className("search_result_item")); var stTime = new Date().getTime(); while (new Date().getTime() < stTime + 15000) { Thread.sleep(500); ((JavascriptExecutor)driver).executeScript("window.scrollTo(0, document.body.scrollHeight)", el); for (WebElement element:el) { System.out.println(++rank+". "); System.out.print(element.findElement(By.tagName("img")).getAttribute("src")); System.out.print("|"+element.findElement(By.className("brand")).getText()); System.out.print("|"+element.findElement(By.className("name")).getText()); System.out.print("|"+element.findElement(By.className("translated_name")).getText()); System.out.print("|"+element.findElement(By.className("amount")).getText()); System.out.print("|"+element.findElement(By.className("desc")).getText()); System.out.print("|"+element.findElement(By.className("express_mark")).getText()); System.out.println(); } } } finally { driver.close(); driver.quit(); } } }
报错信息
onError java.net.SocketException: Connection reset at java.base/sun.nio.ch.SocketChannelImpl.throwConnectionReset(SocketChannelImpl.java:394) at java.base/sun.nio.ch.SocketChannelImpl.read(SocketChannelImpl.java:426) at io.netty.buffer.PooledByteBuf.setBytes(PooledByteBuf.java:258) at io.netty.buffer.AbstractByteBuf.writeBytes(AbstractByteBuf.java:1132) at io.netty.channel.socket.nio.NioSocketChannel.doReadBytes(NioSocketChannel.java:357) at io.netty.channel.nio.AbstractNioByteChannel$NioByteUnsafe.read(AbstractNioByteChannel.java:151) at io.netty.channel.nio.NioEventLoop.processSelectedKey(NioEventLoop.java:722) at io.netty.channel.nio.NioEventLoop.processSelectedKeysOptimized(NioEventLoop.java:658) at io.netty.channel.nio.NioEventLoop.processSelectedKeys(NioEventLoop.java:584) at io.netty.channel.nio.NioEventLoop.run(NioEventLoop.java:496) at io.netty.util.concurrent.SingleThreadEventExecutor$4.run(SingleThreadEventExecutor.java:997) at io.netty.util.internal.ThreadExecutorMap$2.run(ThreadExecutorMap.java:74) at io.netty.util.concurrent.FastThreadLocalRunnable.run(FastThreadLocalRunnable.java:30) at java.base/java.lang.Thread.run(Thread.java:833)
问题根因
- 元素引用失效:商品元素列表在循环外只获取了一次,无限滚动触发DOM更新后,旧的WebElement引用会失效,且
executeScript调用时传入了完全未使用的el参数,额外增加驱动与浏览器的通信开销。 - 反爬拦截:旧版headless模式自动化特征明显,加上固定500ms高频滚动、重复全量遍历DOM的操作,被Kream反爬系统识别后主动断开连接,触发Connection reset。
- 逻辑错误:注释标注滚动30秒,实际代码仅设置15秒滚动时长;每次滚动后全量遍历所有历史商品,重复输出内容的同时额外增加页面负担。
- 资源释放错误:
driver.quit()本身会关闭所有浏览器窗口并释放驱动进程,额外调用driver.close()容易触发通信中断。
修复方案
- 调整Chrome启动参数,隐藏自动化特征
ChromeOptions options = new ChromeOptions(); // 用新版headless模式,和普通浏览器特征一致,避免被检测 options.addArguments("--headless=new"); options.addArguments("--disable-blink-features=AutomationControlled"); // 替换成和当前浏览器版本匹配的UA options.addArguments("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/103.0.5060.114 Safari/537.36"); options.addArguments("--no-sandbox"); options.addArguments("--disable-dev-shm-usage"); // 改用EAGER加载策略,DOM就绪即可操作,不用等所有静态资源加载完成,减少连接超时概率 options.setPageLoadStrategy(PageLoadStrategy.EAGER);
- 修正爬取逻辑
- 把元素列表获取逻辑移到循环内部,每次滚动等待新内容加载完成后再重新拉取元素,避免引用失效
- 删除
executeScript中多余的el传参 - 增加去重逻辑,用HashSet记录已经爬取过的商品图片地址/商品链接,跳过已经爬过的内容,避免重复遍历
- 将固定500ms滚动间隔调整为1000-2000ms的随机间隔,模拟真人浏览行为
- 修正资源释放逻辑
finally块中仅保留driver.quit()即可,删除多余的driver.close()调用。 - 增加容错处理
遍历单个商品元素时捕获StaleElementReferenceException异常,遇到元素失效直接跳过当前项,不要让单个元素错误导致整个程序终止。
内容的提问来源于stack exchange,提问作者hyunseo
相关产品推荐
相关产品推荐

