Other issues

Updated at:

This topic lists other common issues that may occur when you use MaxCompute Spark and provides solutions.

Project self-check

Important

Before you submit a job, for more information, see LogView and check the following items.

  • Check pom.xml : The scope of the spark-xxxx_${scala.binary.version} dependency must be set to provided.

    <dependency>
        <groupId>org.apache.spark</groupId>
        <artifactId>spark-core_${scala.binary.version}</artifactId>
        <version>${spark.version}</version>
        <scope>provided</scope>
    </dependency>
  • Hard-coded spark.master in the main class :

    If you submit the job in yarn-cluster mode, an error occurs if the code contains the local[N] configuration.

    val spark = SparkSession
          .builder()
          .appName("SparkPi")
          // If you submit the job in yarn-cluster mode, do not set the following configuration in the code.
          // .config("spark.master", "local[4]")
          .getOrCreate()
  • The main Scala class must be an object, not a class :

    When you create a file in IntelliJ IDEA, make sure to use an object. Otherwise, the main function cannot be loaded.

    object SparkPi { // This must be an object. Some users might accidentally write "class" when creating a file in IntelliJ IDEA, which prevents the main function from loading.
      def main(args: Array[String]) {
        val spark = SparkSession
          .builder()
          .appName("SparkPi")
          .getOrCreate()
  • Do not hard-code configurations in your code :

    During local testing, developers often hard-code MaxCompute configurations. However, some configurations do not take effect if they are in the code. When you submit a job in yarn-cluster mode, write all configuration items in the spark-defaults.conf file.

    val spark = SparkSession
          .builder()
          .appName("SparkPi")
          .config("key1", "value1")
          .config("key2", "value2")
          .config("key3", "value3")
          ...  // Developers often hard-code MaxCompute configurations in the code during local testing, but some configurations do not take effect.
          .getOrCreate()
    
    // We strongly recommend that you write all configuration items in the spark-defaults.conf file when you submit a job in yarn-cluster mode.

User signature does not match

  • Error message : Invalid signature value - User signature does not match

  • Solution : This error usually occurs because the spark.hadoop.odps.access.id and spark.hadoop.odps.access.key in the spark-defaults.conf file are incorrect. Go to the AccessKey Management page in the Alibaba Cloud console to confirm your AccessKey pair. Make sure no errors occurred during the copy and paste process.

You have NO privilege 'odps:CreateResource'

  • Error message

    com.aliyun.odps.OdpsException: ODPS-0420095:
    Access Denied - Authorization Failed [4019], You have NO privilege 'odps:CreateResource' on {acs:odps:*:projects/*}
  • Solution

    This error usually occurs because the Alibaba Cloud account or RAM user that corresponds to the AccessKey used to submit the Spark job does not have the required permissions. Contact the project owner to grant the required permissions:

    GRANT CreateInstance, CreateResource, ReadResource ON PROJECT <projectName> TO USER <userId>;

    For more information about the authorization syntax, see MaxCompute permissions.

The task is not in release range: CUPID

  • Error message : The task is not in release range: CUPID

  • Solution : This error usually occurs because MaxCompute Spark is not enabled in the current region. Contact technical support to confirm whether the Spark service is available in the region.

No space left on device

  • Error message : No space left on device

  • Solution : Spark uses a network disk as local storage. The driver and each executor have their own network disk. Shuffle data and data that overflows from the BlockManager are stored on the network disk. The disk size is controlled by the spark.hadoop.odps.cupid.disk.driver.device_size parameter. The default size is 20 GB and the maximum size is 100 GB. If this error persists after you increase the disk size to 100 GB, you must analyze the cause:

    • Data skew. During the shuffle or cache process, data is concentrated in specific blocks.

    • Decrease the concurrency of a single executor (spark.executor.cores) and increase the number of executors (spark.executor.instances).

ClassNotFound errors

  • Error message : java.lang.ClassNotFoundException: xxxx.xxx.xxxxx

  • Solution :

    • Run jar -tf <job_jar_package> | grep <class_name> to check whether the submitted JAR package contains the class definition.

    • Check whether the dependencies in the pom.xml file are correct. Make sure that you submit the JAR package that contains the shaded keyword. This keyword indicates that all project dependencies are included in the package.

    • When you submit a Spark job, make sure to submit the package that contains the shaded keyword. This keyword indicates that all project dependencies are included in the package.

JAR package version conflict errors

  • Error message : User class threw exception: java.lang.NoSuchMethodError

  • Solution :

    • In the $SPARK_HOME/jars path, find the JAR package where the abnormal class is located. Run the grep <abnormal_class_name> $SPARK_HOME/jars/*.jar command to locate the coordinates and version of the third-party library.

    • In the root directory of the Spark job project, view all dependencies: mvn dependency:tree

    • After you find the conflicting dependency, use Maven dependency exclusions to exclude the conflicting package. Then, recompile and submit the job.

Shutdown hook called before final status was reported

  • Error message : The job fails soon after it is submitted with the following error: App Status: SUCCEEDED, diagnostics: Shutdown hook called before final status was reported.

  • Solution : This error occurs because the user main function submitted to the cluster exits without requesting cluster resources from the ApplicationMaster.

    • You can create a SparkContext.

    • You can set spark.master to local in the code.

Garbled characters

  • Error message : Garbled characters appear when Chinese characters are printed in a Spark job.

  • Solution

    You can add the following configurations:

    --conf spark.executor.extraJavaOptions="-Dfile.encoding=UTF-8 -Dsun.jnu.encoding=UTF-8"
    --conf spark.driver.extraJavaOptions="-Dfile.encoding=UTF-8 -Dsun.jnu.encoding=UTF-8"

    For a PySpark job, you must also set the following parameters:

    spark.yarn.appMasterEnv.PYTHONIOENCODING=utf8
    spark.executorEnv.PYTHONIOENCODING=utf8

    In addition, you can add the following code at the beginning of the Python script:

    # -*- coding: utf-8 -*-
    import sys
    reload(sys)
    sys.setdefaultencoding('utf-8')

Table or view not found

  • Error message : Table or view not found: xxx

  • Solution :

    First, check whether the table that caused the error exists in the project. If the table exists, check whether the Hive catalog configuration is enabled. If it is, remove the Hive configuration. The following example shows a common PySpark code snippet that causes this error:

    # Incorrect code
    spark = SparkSession.builder.appName(app_name).enableHiveSupport().getOrCreate()
    
    # Correct code: Remove enableHiveSupport
    spark = SparkSession.builder.appName(app_name).getOrCreate()

com.aliyun.odps.cupid.CupidException: Row Not Found

  • Error message : com.aliyun.odps.cupid.CupidException: Row Not Found

  • Solution : This error usually occurs because the project name configured in your code is different from the name of the project where the job runs. The spark.hadoop.odps.project.name parameter must be set to the name of the project where the Spark job runs, not the name of the project where the accessed table is located.

How to reference a JAR package as a resource

You can use spark.hadoop.odps.cupid.resources to upload a JAR package to a project as a resource.

For example: spark.hadoop.odps.cupid.resources=projectname.xx0.jar,projectname.xx1.jar. This method allows resource sharing across multiple projects. You only need to configure the required permissions.

How to adjust the read concurrency for a MaxCompute table

You can configure the split size by setting the spark.hadoop.odps.input.split.size parameter. The unit is MB. The default value is 256. If the data volume of the table is small, you can also read the data and then use repartition to increase the subsequent concurrency.

How to use the MaxCompute Java SDK in code

You can confirm that the cupid-sdk dependency is added with the provided scope. Then, you can obtain the Odps object by calling CupidSession.get().odps().

How to kill a Spark task

You cannot kill a Spark task by pressing Ctrl+C on the Spark client or on the task submission page in the DataWorks console. You can use one of the following methods to kill a running Spark task:

  • You can run kill <instanceId>; on the MaxCompute client.

  • You can click Stop on the DataWorks page.

com.aliyun.odps.cupid.ExceedMaxOpenFileException

  • Error message : An error occurs when you insert data into a dynamic partition.

    com.aliyun.odps.cupid.ExceedMaxOpenFileException: NumPartitions: 1024, maxPartSpecNum: 1000, openFiles: 98, maxOpenFiles: 100000, exceeded Max limit
  • Solution

    This error occurs in the following two cases:

    • This error occurs if NumPartitions × openFiles > maxOpenFiles.

      The default value of maxOpenFiles is controlled by the spark.hadoop.odps.cupid.open.file.max parameter. You can adjust this parameter to a value greater than NumPartitions × openFiles.

    • This error occurs if openFiles > maxPartSpecNum.

      The default value of maxPartSpecNum is controlled by the spark.hadoop.odps.cupid.partition.spec.max parameter. You can adjust this parameter to a value greater than openFiles.

Job execution error: ***.so: cannot open shared object file: No such file or directory

  • Error message : The job execution fails with the error ***.so: cannot open shared object file: No such file or directory

  • Solution

    • MaxCompute Spark client :

      You can download the required dependency file from the public network. When you submit the job, use the --files /path/to/[library_name] parameter to load the dependency file to the working directories of the driver and executors.

    • DataWorks Spark node :

      You can download the required dependency file from the public network. You can add the dependency as a resource in DataWorks by creating a MaxCompute resource. Then, you can add the spark.hadoop.odps.cupid.resources parameter when you submit the job.

    The uploaded dependency resource has a prefix that is the project name. You must rename the uploaded resource to remove the project name prefix so that the dependency can be correctly loaded.