You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spark 2 + Java 8 ANOVA统计计算报错排查请求

Fixing the org.apache.spark.sql.AnalysisException: cannot resolve 'c.sum(valueSq))' Error

Okay, let's break down this error you're hitting. The core issue here is a syntax mistake in your Spark SQL query string—specifically, an extra closing parenthesis that's mangling the column reference.

What's Happening

The error message points to 'c.sum(valueSq))'—notice the extra ) at the end of the column name. When Spark parses your query, it's trying to resolve a column named sum(valueSq)) (with that trailing parenthesis) from the alias c, but that column doesn't exist. This is almost certainly a typo introduced when converting the original Scala code to Java.

Where to Look

Check the SQL string you're using to create the joined DataFrame. You'll find a line where you've written something like:

c.sum(valueSq)) AS groupSumSq

Instead of the correct:

c.sum(valueSq) AS groupSumSq

That extra closing parenthesis is the culprit. It's likely a copy-paste error from the Scala code, where expression syntax might have made the extra bracket less obvious during conversion.

How to Fix It

Locate the SQL query in your Java code where you reference c.sum(valueSq) and remove the extra trailing ). For example, if your query snippet looks like this:

String anovaQuery = "SELECT " +
    "a.groupKey, " +
    "b.totalSum, " +
    "c.sum(value) AS groupSum, " +
    "c.sum(valueSq)) AS groupSumSq " + // <-- Extra ) here!
    "FROM ... JOIN ...";

Change it to:

String anovaQuery = "SELECT " +
    "a.groupKey, " +
    "b.totalSum, " +
    "c.sum(value) AS groupSum, " +
    "c.sum(valueSq) AS groupSumSq " + // Fixed: no extra )
    "FROM ... JOIN ...";

Pro Tip

To avoid these kinds of typos in the future, consider using Spark's typed DataFrame API (with selectExpr or column objects) instead of raw SQL strings. For example, using column references would make this syntax error impossible:

import static org.apache.spark.sql.functions.*;

Dataset<Row> joined = df1.join(df2, ...)
    .join(df3, ...)
    .select(
        col("a.groupKey"),
        col("b.totalSum"),
        sum(col("c.value")).alias("groupSum"),
        sum(col("c.valueSq")).alias("groupSumSq") // No chance for extra parentheses here
    );

内容的提问来源于stack exchange,提问作者F. Aydemir

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 03:49:09