
Enables unified data lake ingestion and management by integrating structured data, unstructured AI data sources, and models

Enables unified data lake ingestion and management by integrating structured data, unstructured AI data sources, and models

A fully managed, batch-stream unified TCIceberg table format. It is fully compatible with Apache Iceberg while extending capabilities for streaming lakehouse scenarios

A fully managed, batch-stream unified TCIceberg table format. It is fully compatible with Apache Iceberg while extending capabilities for streaming lakehouse scenarios

Built-in data catalog for the unified management of multimodal data, with robust support for integrating diverse external data sources

Built-in data catalog for the unified management of multimodal data, with robust support for integrating diverse external data sources

An RBAC-based unified permission model with a standardized access layer, delivering comprehensive access governance across the entire data lifecycle

An RBAC-based unified permission model with a standardized access layer, delivering comprehensive access governance across the entire data lifecycle

Built-in intelligent automation for small file compaction, snapshot cleanup, and overall data lifecycle management

Built-in intelligent automation for small file compaction, snapshot cleanup, and overall data lifecycle management

Open integration with upper-layer compute engines, supporting Tencent Cloud services (e.g., EMR, DLC, TCHouse) and mainstream open-source ecosystems like Apache Spark and Flink

Open integration with upper-layer compute engines, supporting Tencent Cloud services (e.g., EMR, DLC, TCHouse) and mainstream open-source ecosystems like Apache Spark and Flink

Enables unified data lake ingestion and management by integrating structured data, unstructured AI data sources, and models

A fully managed, batch-stream unified TCIceberg table format. It is fully compatible with Apache Iceberg while extending capabilities for streaming lakehouse scenarios

Built-in data catalog for the unified management of multimodal data, with robust support for integrating diverse external data sources

An RBAC-based unified permission model with a standardized access layer, delivering comprehensive access governance across the entire data lifecycle

Built-in intelligent automation for small file compaction, snapshot cleanup, and overall data lifecycle management

Open integration with upper-layer compute engines, supporting Tencent Cloud services (e.g., EMR, DLC, TCHouse) and mainstream open-source ecosystems like Apache Spark and Flink



Frequently
asked questions
The Multimodal Intelligent Data Lake service (TCLake) is a next-generation open, intelligent, and converged AI data lake foundation launched by Tencent Cloud. It provides unified management covering both structured and unstructured data. With built-in features including a multimodal unified data catalog, a batch-stream unified table format, intelligent data management, and data acceleration services, it seamlessly integrates with Tencent Cloud services and mainstream open-source Data+AI ecosystem engines. This empowers enterprises to efficiently build a unified and cost-effective data lake infrastructure for the AI era.
The metadata for Lakehouse, Volume, and Model catalogs is natively stored within the Tencent Cloud TCLake service. This includes metadata services compatible with Hive Metastore, as well as Volume-type data catalogs designed for the unified management of unstructured data. External data catalogs, on the other hand, establish connections with external data sources (such as MySQL) via JDBC or similar methods to retrieve metadata information from those sources in real time.
Currently, Tencent Cloud DLC and TCHouse-D are supported. More data sources are actively being added.
The Multimodal Intelligent Data Lake (TCLake) service is currently in closed beta. This phase is open to invited users only, but you may fill out an application form to request access.
Frequently
asked questions
The Multimodal Intelligent Data Lake service (TCLake) is a next-generation open, intelligent, and converged AI data lake foundation launched by Tencent Cloud. It provides unified management covering both structured and unstructured data. With built-in features including a multimodal unified data catalog, a batch-stream unified table format, intelligent data management, and data acceleration services, it seamlessly integrates with Tencent Cloud services and mainstream open-source Data+AI ecosystem engines. This empowers enterprises to efficiently build a unified and cost-effective data lake infrastructure for the AI era.
The metadata for Lakehouse, Volume, and Model catalogs is natively stored within the Tencent Cloud TCLake service. This includes metadata services compatible with Hive Metastore, as well as Volume-type data catalogs designed for the unified management of unstructured data. External data catalogs, on the other hand, establish connections with external data sources (such as MySQL) via JDBC or similar methods to retrieve metadata information from those sources in real time.
Currently, Tencent Cloud DLC and TCHouse-D are supported. More data sources are actively being added.
The Multimodal Intelligent Data Lake (TCLake) service is currently in closed beta. This phase is open to invited users only, but you may fill out an application form to request access.