Other issues
This topic lists other common issues that may occur when you use MaxCompute Spark and provides solutions.
Project self-check
Before you submit a job, for more information, see LogView and check the following items.
Check pom.xml : The
scopeof thespark-xxxx_${scala.binary.version}dependency must be set toprovided.<dependency> <groupId>org.apache.spark</groupId> <artifactId>spark-core_${scala.binary.version}</artifactId> <version>${spark.version}</version> <scope>provided</scope> </dependency>Hard-coded
spark.masterin the main class :If you submit the job in
yarn-clustermode, an error occurs if the code contains thelocal[N]configuration.val spark = SparkSession .builder() .appName("SparkPi") // If you submit the job in yarn-cluster mode, do not set the following configuration in the code. // .config("spark.master", "local[4]") .getOrCreate()The main Scala class must be an object, not a class :
When you create a file in IntelliJ IDEA, make sure to use an
object. Otherwise, themainfunction cannot be loaded.object SparkPi { // This must be an object. Some users might accidentally write "class" when creating a file in IntelliJ IDEA, which prevents the main function from loading. def main(args: Array[String]) { val spark = SparkSession .builder() .appName("SparkPi") .getOrCreate()Do not hard-code configurations in your code :
During local testing, developers often hard-code MaxCompute configurations. However, some configurations do not take effect if they are in the code. When you submit a job in
yarn-clustermode, write all configuration items in thespark-defaults.conffile.val spark = SparkSession .builder() .appName("SparkPi") .config("key1", "value1") .config("key2", "value2") .config("key3", "value3") ... // Developers often hard-code MaxCompute configurations in the code during local testing, but some configurations do not take effect. .getOrCreate() // We strongly recommend that you write all configuration items in the spark-defaults.conf file when you submit a job in yarn-cluster mode.
User signature does not match
Error message :
Invalid signature value - User signature does not matchSolution : This error usually occurs because the
spark.hadoop.odps.access.idandspark.hadoop.odps.access.keyin thespark-defaults.conffile are incorrect. Go to the AccessKey Management page in the Alibaba Cloud console to confirm your AccessKey pair. Make sure no errors occurred during the copy and paste process.
You have NO privilege 'odps:CreateResource'
Error message
com.aliyun.odps.OdpsException: ODPS-0420095: Access Denied - Authorization Failed [4019], You have NO privilege 'odps:CreateResource' on {acs:odps:*:projects/*}Solution
This error usually occurs because the Alibaba Cloud account or RAM user that corresponds to the AccessKey used to submit the Spark job does not have the required permissions. Contact the project owner to grant the required permissions:
GRANT CreateInstance, CreateResource, ReadResource ON PROJECT <projectName> TO USER <userId>;For more information about the authorization syntax, see MaxCompute permissions.
The task is not in release range: CUPID
Error message :
The task is not in release range: CUPIDSolution : This error usually occurs because MaxCompute Spark is not enabled in the current region. Contact technical support to confirm whether the Spark service is available in the region.
No space left on device
Error message :
No space left on deviceSolution : Spark uses a network disk as local storage. The driver and each executor have their own network disk. Shuffle data and data that overflows from the BlockManager are stored on the network disk. The disk size is controlled by the
spark.hadoop.odps.cupid.disk.driver.device_sizeparameter. The default size is 20 GB and the maximum size is 100 GB. If this error persists after you increase the disk size to 100 GB, you must analyze the cause:Data skew. During the shuffle or cache process, data is concentrated in specific blocks.
Decrease the concurrency of a single executor (
spark.executor.cores) and increase the number of executors (spark.executor.instances).
ClassNotFound errors
Error message :
java.lang.ClassNotFoundException: xxxx.xxx.xxxxxSolution :
Run
jar -tf <job_jar_package> | grep <class_name>to check whether the submitted JAR package contains the class definition.Check whether the dependencies in the
pom.xmlfile are correct. Make sure that you submit the JAR package that contains the shaded keyword. This keyword indicates that all project dependencies are included in the package.When you submit a Spark job, make sure to submit the package that contains the
shadedkeyword. This keyword indicates that all project dependencies are included in the package.
JAR package version conflict errors
Error message :
User class threw exception: java.lang.NoSuchMethodErrorSolution :
In the
$SPARK_HOME/jarspath, find the JAR package where the abnormal class is located. Run thegrep <abnormal_class_name> $SPARK_HOME/jars/*.jarcommand to locate the coordinates and version of the third-party library.In the root directory of the Spark job project, view all dependencies:
mvn dependency:treeAfter you find the conflicting dependency, use Maven dependency exclusions to exclude the conflicting package. Then, recompile and submit the job.
Shutdown hook called before final status was reported
Error message : The job fails soon after it is submitted with the following error:
App Status: SUCCEEDED, diagnostics: Shutdown hook called before final status was reported.Solution : This error occurs because the user main function submitted to the cluster exits without requesting cluster resources from the ApplicationMaster.
You can create a SparkContext.
You can set
spark.mastertolocalin the code.
Garbled characters
Error message : Garbled characters appear when Chinese characters are printed in a Spark job.
Solution
You can add the following configurations:
--conf spark.executor.extraJavaOptions="-Dfile.encoding=UTF-8 -Dsun.jnu.encoding=UTF-8" --conf spark.driver.extraJavaOptions="-Dfile.encoding=UTF-8 -Dsun.jnu.encoding=UTF-8"For a PySpark job, you must also set the following parameters:
spark.yarn.appMasterEnv.PYTHONIOENCODING=utf8 spark.executorEnv.PYTHONIOENCODING=utf8In addition, you can add the following code at the beginning of the Python script:
# -*- coding: utf-8 -*- import sys reload(sys) sys.setdefaultencoding('utf-8')
Table or view not found
Error message :
Table or view not found: xxxSolution :
First, check whether the table that caused the error exists in the project. If the table exists, check whether the Hive catalog configuration is enabled. If it is, remove the Hive configuration. The following example shows a common PySpark code snippet that causes this error:
# Incorrect code spark = SparkSession.builder.appName(app_name).enableHiveSupport().getOrCreate() # Correct code: Remove enableHiveSupport spark = SparkSession.builder.appName(app_name).getOrCreate()
com.aliyun.odps.cupid.CupidException: Row Not Found
Error message :
com.aliyun.odps.cupid.CupidException: Row Not FoundSolution : This error usually occurs because the project name configured in your code is different from the name of the project where the job runs. The
spark.hadoop.odps.project.nameparameter must be set to the name of the project where the Spark job runs, not the name of the project where the accessed table is located.
How to reference a JAR package as a resource
You can use spark.hadoop.odps.cupid.resources to upload a JAR package to a project as a resource.
For example: spark.hadoop.odps.cupid.resources=projectname.xx0.jar,projectname.xx1.jar. This method allows resource sharing across multiple projects. You only need to configure the required permissions.
How to adjust the read concurrency for a MaxCompute table
You can configure the split size by setting the spark.hadoop.odps.input.split.size parameter. The unit is MB. The default value is 256. If the data volume of the table is small, you can also read the data and then use repartition to increase the subsequent concurrency.
How to use the MaxCompute Java SDK in code
You can confirm that the cupid-sdk dependency is added with the provided scope. Then, you can obtain the Odps object by calling CupidSession.get().odps().
How to kill a Spark task
You cannot kill a Spark task by pressing Ctrl+C on the Spark client or on the task submission page in the DataWorks console. You can use one of the following methods to kill a running Spark task:
You can run
kill <instanceId>;on the MaxCompute client.You can click Stop on the DataWorks page.
com.aliyun.odps.cupid.ExceedMaxOpenFileException
Error message : An error occurs when you insert data into a dynamic partition.
com.aliyun.odps.cupid.ExceedMaxOpenFileException: NumPartitions: 1024, maxPartSpecNum: 1000, openFiles: 98, maxOpenFiles: 100000, exceeded Max limitSolution
This error occurs in the following two cases:
This error occurs if
NumPartitions × openFiles > maxOpenFiles.The default value of
maxOpenFilesis controlled by thespark.hadoop.odps.cupid.open.file.maxparameter. You can adjust this parameter to a value greater thanNumPartitions × openFiles.This error occurs if
openFiles > maxPartSpecNum.The default value of
maxPartSpecNumis controlled by thespark.hadoop.odps.cupid.partition.spec.maxparameter. You can adjust this parameter to a value greater thanopenFiles.
Job execution error: ***.so: cannot open shared object file: No such file or directory
Error message : The job execution fails with the error
***.so: cannot open shared object file: No such file or directorySolution
MaxCompute Spark client :
You can download the required dependency file from the public network. When you submit the job, use the
--files /path/to/[library_name]parameter to load the dependency file to the working directories of the driver and executors.DataWorks Spark node :
You can download the required dependency file from the public network. You can add the dependency as a resource in DataWorks by creating a MaxCompute resource. Then, you can add the
spark.hadoop.odps.cupid.resourcesparameter when you submit the job.
The uploaded dependency resource has a prefix that is the project name. You must rename the uploaded resource to remove the project name prefix so that the dependency can be correctly loaded.