All Products
Search
Document Center

MaxCompute:Developing Java UDFs

Last Updated:Aug 14, 2026

If MaxCompute built-in functions do not meet your requirements, you can create a user-defined function (UDF) in Java by using a development tool such as IntelliJ IDEA (Maven) or MaxCompute Studio, and then call the UDF in MaxCompute SQL statements.

Limits

  • Access the Internet by using UDFs

    By default, MaxCompute does not allow you to access the Internet by using UDFs. If you want to access the Internet by using UDFs, fill in the network connection application form based on your business requirements and submit the application. After the application is approved, the MaxCompute technical support team will contact you and help you establish network connections. For more information about how to fill in the network connection application form, see Network connection process.

  • Access a VPC by using UDFs

    By default, MaxCompute does not allow you to access resources in VPCs by using UDFs. To use UDFs to access resources in a VPC, you must establish a network connection between MaxCompute and the VPC. For more information about related operations, see Access VPC resources from a UDF.

  • Read table data by using UDFs, UDAFs, or UDTFs

    You cannot use UDFs, UDAFs, or UDTFs to read data from the following types of tables:

    • Table on which schema evolution is performed

    • Table that contains complex data types

    • Table that contains JSON data types

    • Transactional table

Usage notes

Before you write a Java UDF, familiarize yourself with the UDF code structure and the data type mappings between Java and MaxCompute. For more information, see Appendix: Data types.

When you write a Java UDF, take note of the following:

  • Avoid including classes that have the same name but different logic in different UDF JAR files. For example, if UDF1 and UDF2 correspond to udf1.jar and udf2.jar respectively, and both JAR files contain com.aliyun.UserFunction.class with different logic, calling both UDFs in the same SQL statement causes MaxCompute to randomly load one of the classes. This can lead to unexpected results or a compilation failure.

  • In a Java UDF, input parameters and return values must use object types (such as String and Long), not primitive types (such as int and long).

  • NULL values in SQL map to NULL in Java. Java primitive types cannot represent NULL values and are therefore not allowed.

UDF development workflow

UDF development involves several steps: preparing the environment, writing UDF code, uploading the JAR file, registering the UDF, and debugging. The following sections demonstrate this workflow with MaxCompute Studio, DataWorks, and odpscmd.

Use MaxCompute Studio

The following example shows how to develop and call a Java UDF that converts characters to lowercase by using MaxCompute Studio.

  1. Prepare the environment.

    Before you can develop and debug a UDF in MaxCompute Studio, install MaxCompute Studio and connect it to a MaxCompute project. For more information, see the following topics:

    1. Install MaxCompute Studio

    2. Connect to a MaxCompute project

    3. Create a MaxCompute Java module

  2. Write UDF code.

    1. In the Project explorer, right-click the source code directory of the module (src > main > java) and select New > MaxCompute Java.

    2. In the Create new MaxCompute java class dialog box, click UDF, enter a class name in the Name field, and then press Enter.

      Name specifies the name of the MaxCompute Java class to create. If you have not created a package, enter packagename.classname to automatically create one. In this example, the class is named Lower.

    3. Write the UDF code in the code editor. For example:

      package com.aliyun.odps.udf.example;
      import com.aliyun.odps.udf.UDF;
      public final class Lower extends UDF {
          public String evaluate(String s) {
              if (s == null) { 
                 return null; 
              }
                 return s.toLowerCase();
          }
      }
      Note

      To locally debug the Java UDF, see Develop and debug UDFs.

  3. Upload and register the UDF.

    Right-click the UDF Java file and select Deploy to server.... In the Package a jar, submit resource and register function dialog box, configure the parameters and click OK.

    • MaxCompute project: The MaxCompute project to which the UDF belongs. Because the UDF is written in the connected MaxCompute project, you can use the default value.

    • Resource file: The path of the resource file that the UDF depends on. You can use the default value.

    • Resource name: The resource that the UDF depends on. You can use the default value.

    • Function name: The name used to call the UDF in SQL statements. For example, Lower_test.

  4. Debug the UDF.

    In the left-side navigation pane, click Project Explore. Right-click the destination MaxCompute project and select Open Console. In the console, enter the SQL statement that calls the UDF and press Enter. For example:

    select lower_test('ABC');

    The following result is returned.

    +-----+
    | _c0 |
    +-----+
    | abc |
    +-----+

Use DataWorks

  1. Prepare the environment.

    Before you can develop and debug a UDF in DataWorks, activate DataWorks and bind a MaxCompute project to it. For more information, see Use DataWorks.

  2. Write UDF code.

    You can write UDF code in any Java development tool and package it as a JAR file. For example:

    package com.aliyun.odps.udf.example;
    import com.aliyun.odps.udf.UDF;
    public final class Lower extends UDF {
        public String evaluate(String s) {
            if (s == null) { 
               return null; 
            }
               return s.toLowerCase();
        }
    }
  3. Upload and register the UDF.

    You can upload the packaged code to DataWorks and register the UDF. For more information, see the following topics:

    1. Create and use MaxCompute resources

    2. Create and use a MaxCompute function

  4. Debug the UDF.

    After you register the UDF, create an ODPS SQL node and run a SQL statement in the node to debug the UDF. For more information about how to create an ODPS SQL node, see Create an ODPS SQL node. Example:

    select lower_test('ABC');

Use odpscmd

  1. Prepare the environment.

    To develop and debug a UDF by using odpscmd, install the client and configure its connection to a MaxCompute project. For more information, see Use the MaxCompute client (odpscmd).

  2. Write UDF code.

    You can write UDF code in any Java development tool and package it as a JAR file. For example:

    package com.aliyun.odps.udf.example;
    import com.aliyun.odps.udf.UDF;
    public final class Lower extends UDF {
        public String evaluate(String s) {
            if (s == null) { 
               return null; 
            }
               return s.toLowerCase();
        }
    }
  3. Upload and register the UDF.

    You can upload the packaged code by using odpscmd and register the UDF. For more information, see:

    1. ADD JAR

    2. CREATE FUNCTION

  4. Debug the UDF.

    After you register the UDF, write and run a SQL statement to debug it. Example:

    select lower_test('ABC');

Calling UDFs

After you develop a Java UDF as described in the UDF development workflow, you can call it in MaxCompute SQL. The following methods are available:

  • Use a UDF in a MaxCompute project: The method is similar to that of using built-in functions.

  • Use a UDF across projects: Use a UDF of Project B in Project A. The following statement shows an example: select B:udf_in_other_project(arg0, arg1) as res from table_t;. For more information about cross-project sharing, see Cross-project resource access based on packages.

UDF examples

Appendix: UDF code structure

A Java UDF consists of the following parts:

  • Java package: Optional.

    You can group your Java classes into a package for easier reuse and organization.

  • Inherit the UDF class: Required.

    The required base class is com.aliyun.odps.udf.UDF. If you need other UDF classes or complex data types, add the required classes from the MaxCompute SDK. For example, the class for the STRUCT data type is com.aliyun.odps.data.Struct.

  • @Resolve annotation: Optional.

    The format is @Resolve(<signature>), where signature defines the data types of the input parameters and return value. When you use the STRUCT data type in a UDF, reflection cannot retrieve field names and field types from com.aliyun.odps.data.Struct. In this case, you must use the @Resolve annotation to retrieve them. If you use STRUCT in a UDF, add the @Resolve annotation to the UDF class. The annotation affects only the overloads whose parameters or return values contain com.aliyun.odps.data.Struct. Example: @Resolve("struct<a:string>,string->string"). For a detailed example, see UDF Example: Complex Data Types.

  • Custom Java class: Required.

    This is the unit that organizes your UDF code and defines the variables and methods that implement your business logic.

  • evaluate method: Required.

    Your custom Java class must include a non-static public evaluate method. The data types of its input parameters and return value define the UDF's SQL signature.

    You can implement multiple evaluate methods. When you call the UDF, MaxCompute selects the matching evaluate method based on the argument types.

    When you write a Java UDF, you can use Java types or Java Writable types. For detailed mappings between MaxCompute data types and Java data types, see Appendix: Data types.

  • UDF initialization and cleanup: Optional. You can implement initialization and cleanup using void setup(ExecutionContext ctx) and void close(). The void setup(ExecutionContext ctx) method is called once before the evaluate method and can be used to initialize resources or member objects required for the computation. The void close() method is called once after all evaluate calls are complete and is used for cleanup tasks, such as closing files.

The following examples show two types of UDFs.

  • Use Java types

    // Organize the Java class in the org.alidata.odps.udf.examples package.
    package org.alidata.odps.udf.examples;  
    // Inherit the UDF class.
    import com.aliyun.odps.udf.UDF;         
    // Define a custom Java class.
    public final class Lower extends UDF { 
    // The evaluate method defines the UDF's logic. It takes a String and returns a String.
        public String evaluate(String s) { 
            if (s == null) { 
            return null; 
        } 
            return s.toLowerCase(); 
      } 
    }
  • Use Java Writable types

    // Organize the Java class in the com.aliyun.odps.udf.example package.
    package com.aliyun.odps.udf.example;
    // Add the required classes for the Java Writable type.
    import com.aliyun.odps.io.Text;
    // Inherit the UDF class.
    import com.aliyun.odps.udf.UDF;
    // Define a custom Java class.
    public class MyConcat extends UDF {
      private Text ret = new Text();
    // Define the evaluate method. `Text` specifies the data type of the input parameters, and the `return` value is also a Text object.
      public Text evaluate(Text a, Text b) {
          if (a == null || b == null) {
          return null;
        }
          ret.clear();
          ret.append(a.getBytes(), 0, a.getLength());
          ret.append(b.getBytes(), 0, b.getLength());
          return ret;
      }
    }

MaxCompute also supports UDFs that are developed for its compatible Hive version. For more information, see Hive UDF compatibility.

Appendix: Data types

Data type mappings

To ensure that the data types used in a Java UDF are consistent with MaxCompute data types, use the following mappings.

Note

The data types supported by MaxCompute vary by data type edition. Starting from MaxCompute 2.0, additional data types are available, including complex types such as ARRAY, MAP, and STRUCT. For more information, see Data type editions.

MaxCompute type

Java type

Java Writable type

TINYINT

java.lang.Byte

ByteWritable

SMALLINT

java.lang.Short

ShortWritable

INT

java.lang.Integer

IntWritable

BIGINT

java.lang.Long

LongWritable

FLOAT

java.lang.Float

FloatWritable

DOUBLE

java.lang.Double

DoubleWritable

DECIMAL

java.math.BigDecimal

BigDecimalWritable

BOOLEAN

java.lang.Boolean

BooleanWritable

STRING

java.lang.String

Text

VARCHAR

com.aliyun.odps.data.Varchar

VarcharWritable

BINARY

com.aliyun.odps.data.Binary

BytesWritable

DATE

java.sql.Date

DateWritable

DATETIME

java.util.Date

DatetimeWritable

TIMESTAMP

java.sql.Timestamp

TimestampWritable

INTERVAL_YEAR_MONTH

N/A

IntervalYearMonthWritable

INTERVAL_DAY_TIME

N/A

IntervalDayTimeWritable

ARRAY

java.util.List

N/A

MAP

java.util.Map

N/A

STRUCT

com.aliyun.odps.data.Struct

N/A

The Java byte[] type is not in the list of supported Java types. If you use byte[] for an input parameter or the return value of an evaluate method, an ODPS-0130071 error is reported. To process binary data in a UDF, use the Java types corresponding to the MaxCompute BINARY type: com.aliyun.odps.data.Binary for the standard Java type, or com.aliyun.odps.io.BytesWritable for the Writable type. You can also convert binary data to a Base64-encoded string and use the String type for the return value.

Hive UDF compatibility

If your MaxCompute project uses the 2.0 data type edition, MaxCompute supports Hive-style UDFs. You can directly use Hive UDFs developed for a compatible Hive version.

The compatible Hive version is 2.1.0, which corresponds to Hadoop 2.7.2. If your UDF was compiled against a different Hive or Hadoop version, recompile the UDF JAR file with Hive 2.1.0 or Hadoop 2.7.2.

For a detailed example of using a Hive UDF in MaxCompute, see UDF Example: Hive Compatibility.