You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spark on YARN中通过NodeManager修改配置不生效问题求助

问题分析与解决方案

核心问题排查

你修改__spark_hadoop_conf__.xml后配置未生效,大概率是以下几个原因:

1. 修改了错误的配置文件

从你提供的文件列表来看,__spark_conf__.zip解压后存在__hadoop_conf__目录,Spark容器启动时实际加载的是该目录下的core-site.xml,而非根目录的__spark_hadoop_conf__.xml。你修改的文件并非Spark真正读取的目标文件。

2. 文件权限不符合要求

原配置文件权限是r-x------,但你修改后的新文件权限为rw-rw-r--,YARN容器进程以特定用户运行,可能没有读取该文件的权限,导致配置加载失败。

3. 修改时机太晚

在DefaultContainerExecutor.launchContainer方法中执行文件修改时,YARN可能已经完成了配置的解析或传递,容器启动脚本已经读取了原配置,后续修改文件无法影响已加载的配置。


针对性解决方案

方案一:修改正确的配置文件路径

调整代码中遍历的目标文件路径,定位到__hadoop_conf__目录下的core-site.xml:

@Override
public int launchContainer(ContainerStartContext ctx) throws IOException, ConfigurationException {
    // ... 原有代码
    Log.info("--------------- change start --------------------------------");
    Map<Path, List<String>> localizedResources = ctx.getLocalizedResources();
    Set<Path> paths = localizedResources.keySet();
    for (Path path : paths) {
        String pathName = path.toUri().getPath();
        if (StringUtils.isNotEmpty(pathName) && pathName.contains("__spark_conf__.zip")) {
            // 遍历__hadoop_conf__目录
            File hadoopConfDir = new File(pathName, "__hadoop_conf__");
            if (hadoopConfDir.exists() && hadoopConfDir.isDirectory()) {
                File coreSiteXml = new File(hadoopConfDir, "core-site.xml");
                if (coreSiteXml.exists()) {
                    // 备份原文件
                    File backupFile = new File(hadoopConfDir, "core-site.xml.bak");
                    Files.copy(coreSiteXml.toPath(), backupFile.toPath(), StandardCopyOption.REPLACE_EXISTING);
                    
                    // 解析并修改XML配置
                    DocumentBuilderFactory factory = DocumentBuilderFactory.newInstance();
                    DocumentBuilder builder = factory.newDocumentBuilder();
                    Document doc = builder.parse(coreSiteXml);
                    NodeList properties = doc.getElementsByTagName("property");
                    for (int i = 0; i < properties.getLength(); i++) {
                        Node prop = properties.item(i);
                        Node nameNode = prop.getFirstChild();
                        while (nameNode != null && !nameNode.getNodeName().equals("name")) {
                            nameNode = nameNode.getNextSibling();
                        }
                        if (nameNode != null && "my.target.keyA".equals(nameNode.getTextContent().trim())) {
                            Node valueNode = nameNode.getNextSibling();
                            while (valueNode != null && !valueNode.getNodeName().equals("value")) {
                                valueNode = valueNode.getNextSibling();
                            }
                            if (valueNode != null) {
                                valueNode.setTextContent("MY_TARGET_VALUE_BBB");
                                break;
                            }
                        }
                    }
                    // 写入修改后的XML
                    Transformer transformer = TransformerFactory.newInstance().newTransformer();
                    transformer.setOutputProperty(OutputKeys.INDENT, "yes");
                    transformer.transform(new DOMSource(doc), new StreamResult(coreSiteXml));
                    
                    // 还原原文件权限
                    coreSiteXml.setReadable(true, false);
                    coreSiteXml.setExecutable(true, false);
                    coreSiteXml.setWritable(false, false);
                }
            }
        }
    }
    Log.info("--------------- change end --------------------------------");
    // ... 原有代码
}

方案二:通过环境变量覆盖配置(更简洁)

无需修改XML文件,直接在launchContainer方法中添加环境变量覆盖配置:

@Override
public int launchContainer(ContainerStartContext ctx) throws IOException, ConfigurationException {
    // ... 原有代码
    Map<String, String> env = ctx.getEnvironment();
    // 添加Spark配置覆盖参数
    String sparkConf = env.get("SPARK_CONF");
    sparkConf = (sparkConf == null) ? "" : sparkConf;
    sparkConf += " --conf my.target.keyA=MY_TARGET_VALUE_BBB";
    env.put("SPARK_CONF", sparkConf);
    // ... 原有代码
}

这种方式直接通过Spark启动参数覆盖配置,绕开了文件修改的权限和时机问题,更可靠。


验证步骤

  1. 重新编译YARN代码,替换NodeManager的jar包并重启服务
  2. 提交Spark任务后,进入容器本地目录,检查__hadoop_conf__/core-site.xml中的配置值是否已修改(方案一适用)
  3. 查看Spark任务的日志,搜索my.target.keyA,确认加载的是修改后的值

内容的提问来源于stack exchange,提问作者AppleCEO

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 23:10:22