Spark on YARN中通过NodeManager修改配置不生效问题求助
问题分析与解决方案
核心问题排查
你修改__spark_hadoop_conf__.xml后配置未生效,大概率是以下几个原因:
1. 修改了错误的配置文件
从你提供的文件列表来看,__spark_conf__.zip解压后存在__hadoop_conf__目录,Spark容器启动时实际加载的是该目录下的core-site.xml,而非根目录的__spark_hadoop_conf__.xml。你修改的文件并非Spark真正读取的目标文件。
2. 文件权限不符合要求
原配置文件权限是r-x------,但你修改后的新文件权限为rw-rw-r--,YARN容器进程以特定用户运行,可能没有读取该文件的权限,导致配置加载失败。
3. 修改时机太晚
在DefaultContainerExecutor.launchContainer方法中执行文件修改时,YARN可能已经完成了配置的解析或传递,容器启动脚本已经读取了原配置,后续修改文件无法影响已加载的配置。
针对性解决方案
方案一:修改正确的配置文件路径
调整代码中遍历的目标文件路径,定位到__hadoop_conf__目录下的core-site.xml:
@Override public int launchContainer(ContainerStartContext ctx) throws IOException, ConfigurationException { // ... 原有代码 Log.info("--------------- change start --------------------------------"); Map<Path, List<String>> localizedResources = ctx.getLocalizedResources(); Set<Path> paths = localizedResources.keySet(); for (Path path : paths) { String pathName = path.toUri().getPath(); if (StringUtils.isNotEmpty(pathName) && pathName.contains("__spark_conf__.zip")) { // 遍历__hadoop_conf__目录 File hadoopConfDir = new File(pathName, "__hadoop_conf__"); if (hadoopConfDir.exists() && hadoopConfDir.isDirectory()) { File coreSiteXml = new File(hadoopConfDir, "core-site.xml"); if (coreSiteXml.exists()) { // 备份原文件 File backupFile = new File(hadoopConfDir, "core-site.xml.bak"); Files.copy(coreSiteXml.toPath(), backupFile.toPath(), StandardCopyOption.REPLACE_EXISTING); // 解析并修改XML配置 DocumentBuilderFactory factory = DocumentBuilderFactory.newInstance(); DocumentBuilder builder = factory.newDocumentBuilder(); Document doc = builder.parse(coreSiteXml); NodeList properties = doc.getElementsByTagName("property"); for (int i = 0; i < properties.getLength(); i++) { Node prop = properties.item(i); Node nameNode = prop.getFirstChild(); while (nameNode != null && !nameNode.getNodeName().equals("name")) { nameNode = nameNode.getNextSibling(); } if (nameNode != null && "my.target.keyA".equals(nameNode.getTextContent().trim())) { Node valueNode = nameNode.getNextSibling(); while (valueNode != null && !valueNode.getNodeName().equals("value")) { valueNode = valueNode.getNextSibling(); } if (valueNode != null) { valueNode.setTextContent("MY_TARGET_VALUE_BBB"); break; } } } // 写入修改后的XML Transformer transformer = TransformerFactory.newInstance().newTransformer(); transformer.setOutputProperty(OutputKeys.INDENT, "yes"); transformer.transform(new DOMSource(doc), new StreamResult(coreSiteXml)); // 还原原文件权限 coreSiteXml.setReadable(true, false); coreSiteXml.setExecutable(true, false); coreSiteXml.setWritable(false, false); } } } } Log.info("--------------- change end --------------------------------"); // ... 原有代码 }
方案二:通过环境变量覆盖配置(更简洁)
无需修改XML文件,直接在launchContainer方法中添加环境变量覆盖配置:
@Override public int launchContainer(ContainerStartContext ctx) throws IOException, ConfigurationException { // ... 原有代码 Map<String, String> env = ctx.getEnvironment(); // 添加Spark配置覆盖参数 String sparkConf = env.get("SPARK_CONF"); sparkConf = (sparkConf == null) ? "" : sparkConf; sparkConf += " --conf my.target.keyA=MY_TARGET_VALUE_BBB"; env.put("SPARK_CONF", sparkConf); // ... 原有代码 }
这种方式直接通过Spark启动参数覆盖配置,绕开了文件修改的权限和时机问题,更可靠。
验证步骤
- 重新编译YARN代码,替换NodeManager的jar包并重启服务
- 提交Spark任务后,进入容器本地目录,检查
__hadoop_conf__/core-site.xml中的配置值是否已修改(方案一适用) - 查看Spark任务的日志,搜索
my.target.keyA,确认加载的是修改后的值
内容的提问来源于stack exchange,提问作者AppleCEO
相关产品推荐
相关产品推荐

