Use DLF shared sample datasets

Updated at:

DLF provides TPC-DS sample databases of multiple sizes and a search sample dataset. Use them to validate traditional big data queries and performance, or multimodal scenarios such as image, vector, and full-text search. This topic describes the available datasets and how to create and use a read-only sample catalog.

Select a sample dataset

DLF provides the following types of shared sample datasets:

  • TPC-DS Paimon and TPC-DS Iceberg datasets: Use these datasets for traditional big data queries, analytics, and performance validation.
  • Search sample dataset: This dataset contains images, vectors, and official documentation data for multimodal scenarios such as image, vector, and full-text search.

The following datasets are available:

Sample database name

Sample data description

tpcds_paimon_sf1

TPC-DS 1 GB Paimon table

tpcds_paimon_sf2

TPC-DS 2 GB Paimon table

tpcds_paimon_sf10

TPC-DS 10 GB Paimon table

tpcds_paimon_sf100

TPC-DS 100 GB Paimon table

tpcds_iceberg_sf1

TPC-DS 1 GB Iceberg table

search_samples

Search sample dataset with images, vectors, and official documentation data for image retrieval, vector retrieval, and full-text search scenarios

Create a read-only sample catalog

  1. Log on to the DLF console
  2. In the left navigation pane, click Catalogs.

  3. Click Data Sharing > Shared With Me, locate the data share named dlf_samples, and click Create Catalog.

    The catalog created from a received share is read-only.

  4. Click the Catalogs tab to view the newly created catalog.

    In the catalog list, the status of the new catalog appears as Running.

Use sample datasets

Sample datasets are created and maintained by the DLF team and support multimodal capabilities including image search, visual exploration, and full-text search. For detailed instructions, see Quickly experience DLF multimodal retrieval.