Work with DLF Lance using Python
This guide shows how to use lance-dlf to connect to DLF Catalog, create Lance tables, write data, and verify results.
What it does
lance-dlf provides these core features:
Connects to DLF Catalog
Maps DLF databases to Lance namespaces (logical mapping layer)
Exposes only
type=lance-tableDLF tablesObtains temporary OSS access credentials via DLF's
load_table_tokenAPIConverts OSS temporary credentials into
storage_optionsfor PyLance
PyLance handles the actual data read/write operations:
# Write data
lance.write_dataset(table, location, storage_options=storage_options)
# Read data
lance.dataset(location, storage_options=storage_options)Install lance-dlf
Install lance-dlf from PyPI:
python3 -m pip install lance-dlfConfigure catalog connection
Minimal configuration
CONFIG = {
"uri": "http://<dlf-endpoint>",
"warehouse": "<warehouse>",
"token.provider": "dlf",
"dlf.region": "<region>",
"dlf.access-key-id": "<access-key-id>",
"dlf.access-key-secret": "<access-key-secret>",
"dlf.oss-endpoint": "<oss-endpoint>",
}Configuration parameters
Parameter | Description |
| DLF REST Catalog endpoint, e.g. |
| Warehouse name for the DLF Catalog |
| Use |
| DLF region, e.g. |
| AccessKey ID for DLF access |
| AccessKey Secret for DLF access |
| (Optional) Security token for STS scenarios |
| (Optional) Custom OSS endpoint, e.g. |
Important: Your AccessKey ID and Secret are critical credentials to access Alibaba Cloud resources. Store them securely and never commit real AccessKeys to Git repositories. Read credentials from:
Environment variables
A key management system
Runtime configuration
Connect to DLF Catalog
Importing the lance_dlf module automatically registers the dlf namespace implementation:
import lance_namespace
import lance_dlf # noqa: F401
# Connect to DLF Catalog
ns = lance_namespace.connect("dlf", CONFIG)
# Verify the connection
print(ns.namespace_id())The namespace_id() method returns information about the connected catalog, including the DLF endpoint and warehouse.
View namespaces and tables
Concept mapping
DLF databases map to Lance namespaces.
Query examples
from lance_namespace import (
DescribeNamespaceRequest,
DescribeTableRequest,
ListNamespacesRequest,
ListTablesRequest,
)
DATABASE = "<database>"
TABLE = "<table>"
# List all namespaces
namespaces = ns.list_namespaces(ListNamespacesRequest(id=[]))
print(namespaces)
# Get namespace details
namespace = ns.describe_namespace(DescribeNamespaceRequest(id=[DATABASE]))
print(namespace)
# List all tables in the database
tables = ns.list_tables(ListTablesRequest(id=[DATABASE]))
print(tables)
# Get table details
table = ns.describe_table(DescribeTableRequest(id=[DATABASE, TABLE]))
print(table.location)
print(table.properties)
print(table.storage_options)describe_table return fields
Field | Description |
| Physical storage path for the Lance dataset (typically |
| DLF table schema options, |
| Temporary access credentials for PyLance to read/write OSS |
Important: storage_options contains temporary AccessKey, AccessKey Secret, and Token. Always redact these values in logs.
Create Lance table and write data
Create new table
If the table does not exist, create it by using the low-level API:
Build a DLF schema from the Arrow schema, then create the table metadata in DLF.
Get the Lance table storage location (
location) after the table is created.Build temporary OSS access credentials through
lance-dlf.Write Arrow data to the specified location with PyLance, then notify DLF of the Lance commit.
import lance
import pyarrow as pa
from lance_dlf.api.rest_exception import AlreadyExistsException
from lance_dlf.common.identifier import Identifier
from lance_dlf.schema.schema import Schema as DlfSchema
DATABASE = "default"
TABLE = "test_lance_create_001"
table_id = [DATABASE, TABLE]
# Build test data
data = pa.table({
"f0": pa.array([101, 102, 103], type=pa.int64()),
"f1": pa.array(["create-a", "create-b", "create-c"], type=pa.string()),
})
# Create table and write data
expected = DlfSchema.from_pyarrow_schema(data.schema, options={"type": "lance-table"})
identifier = Identifier(*table_id)
try:
ns._api.create_table(identifier, expected)
except AlreadyExistsException:
response = ns._api.get_table(identifier)
if response.get_schema() != expected:
raise ValueError("The existing table schema does not match the schema of the data to write")
else:
response = ns._api.get_table(identifier)
location = response.get_path()
storage_options = ns._build_storage_options(identifier)
lance.write_dataset(data, location, mode="overwrite", storage_options=storage_options)
ns.notify_lance_commit(table_id)
# Read and verify
dataset = lance.dataset(location, storage_options=storage_options)
result = dataset.to_table()
print(result)Expected output
pyarrow.Table
f0: int64
f1: string
----
f0: [[101,102,103]]
f1: [["create-a","create-b","create-c"]]Write to existing empty table
If an empty type=lance-table table exists in DLF, get the table's storage location and access credentials through describe_table, then write data with PyLance.
import lance
import pyarrow as pa
from lance_namespace import DescribeTableRequest
DATABASE = "default"
TABLE = "test_lance_table"
table_id = [DATABASE, TABLE]
# Get table details
desc = ns.describe_table(DescribeTableRequest(id=table_id))
# Build test data
data = pa.table({
"f0": pa.array([1, 2, 3], type=pa.int64()),
"f1": pa.array(["value-1", "value-2", "value-3"], type=pa.string()),
})
# Write data
lance.write_dataset(
data,
desc.location,
mode="overwrite",
storage_options=desc.storage_options,
)
ns.notify_lance_commit(table_id)
# Read and verify
dataset = lance.dataset(desc.location, storage_options=desc.storage_options)
print(dataset.to_table())Write modes
Mode | Description |
| Overwrites existing data; useful to initialize empty tables or test tables |
| Appends data; the schema must be compatible |
Warning: overwrite mode replaces all existing Lance dataset data. Be careful when using this mode.
Build test data from DLF table schema
Instead of hardcoding column names and types, extract field information from the DLF table schema to build test data.
import pyarrow as pa
from lance_dlf.common.identifier import Identifier
def sample_value(field_type: str, row: int):
"""Generate sample value by field type"""
normalized = field_type.lower()
if "int" in normalized:
return row + 1
if "string" in normalized or "char" in normalized or "varchar" in normalized:
return f"value-{row + 1}"
raise ValueError(f"Unsupported sample field type: {field_type}")
def build_sample_table(ns, database: str, table: str) -> pa.Table:
"""Build sample data from DLF table schema"""
raw_table = ns._api.get_table(Identifier(database, table))
schema = raw_table.get_schema()
fields = schema.fields if schema and schema.fields else []
data = {}
for field in fields:
field_type = str(field.type)
data[field.name] = [sample_value(field_type, row) for row in range(3)]
return pa.table(data)Usage example
desc = ns.describe_table(DescribeTableRequest(id=[DATABASE, TABLE]))
data = build_sample_table(ns, DATABASE, TABLE)
lance.write_dataset(
data,
desc.location,
mode="overwrite",
storage_options=desc.storage_options,
)Complete example
This example demonstrates the full workflow: connecting to DLF, creating a table, writing data, listing tables, and verifying results.
from datetime import datetime
import lance
import lance_namespace
import pyarrow as pa
from lance_dlf.api.rest_exception import AlreadyExistsException
from lance_dlf.common.identifier import Identifier
from lance_dlf.schema.schema import Schema as DlfSchema
from lance_namespace import ListTablesRequest
import lance_dlf # noqa: F401
CONFIG = {
"uri": "http://<dlf-endpoint>",
"warehouse": "<warehouse>",
"token.provider": "dlf",
"dlf.region": "<region>",
"dlf.access-key-id": "<access-key-id>",
"dlf.access-key-secret": "<access-key-secret>",
"dlf.oss-endpoint": "<oss-endpoint>",
}
DATABASE = "default"
def main():
# Connect to DLF Catalog
ns = lance_namespace.connect("dlf", CONFIG)
# Generate unique table name
table_name = "test_lance_create_" + datetime.now().strftime("%Y%m%d_%H%M%S")
table_id = [DATABASE, table_name]
# Build test data
data = pa.table({
"f0": pa.array([101, 102, 103], type=pa.int64()),
"f1": pa.array(["create-a", "create-b", "create-c"], type=pa.string()),
})
# Create table and write data
expected = DlfSchema.from_pyarrow_schema(data.schema, options={"type": "lance-table"})
identifier = Identifier(*table_id)
try:
ns._api.create_table(identifier, expected)
except AlreadyExistsException:
response = ns._api.get_table(identifier)
if response.get_schema() != expected:
raise ValueError("The existing table schema does not match the schema of the data to write")
else:
response = ns._api.get_table(identifier)
location = response.get_path()
storage_options = ns._build_storage_options(identifier)
lance.write_dataset(data, location, mode="overwrite", storage_options=storage_options)
ns.notify_lance_commit(table_id)
print("created:", ".".join(table_id))
print("location:", location)
# List tables and verify
tables = ns.list_tables(ListTablesRequest(id=[DATABASE]))
print("listed_after_create:", table_name in tables.tables)
# Read data and verify
dataset = lance.dataset(location, storage_options=storage_options)
result = dataset.to_table()
print(result)
# Data integrity check
expected = data.to_pylist()
actual = result.to_pylist()
if actual != expected:
raise AssertionError(f"readback mismatch: expected={expected}, actual={actual}")
print("create_write_read: ok")
if __name__ == "__main__":
main()Run the example
python3 main.py