tencent cloud

Tencent Cloud TCLake
TCLake is Tencent Cloud's next-generation Data+AI lake foundation. It unifies structured and unstructured data with a multimodal catalog and a batch-stream unified format, delivering a cost-effective, AI-ready infrastructure for modern enterprise engines.
Features of Tencent Cloud TCLake
Multimodal Data Integration

Enables unified data lake ingestion and management by integrating structured data, unstructured AI data sources, and models

Multimodal Data Integration

Enables unified data lake ingestion and management by integrating structured data, unstructured AI data sources, and models

Batch-Stream Unified Table Format

A fully managed, batch-stream unified TCIceberg table format. It is fully compatible with Apache Iceberg while extending capabilities for streaming lakehouse scenarios

Batch-Stream Unified Table Format

A fully managed, batch-stream unified TCIceberg table format. It is fully compatible with Apache Iceberg while extending capabilities for streaming lakehouse scenarios

Unified Data Catalog

Built-in data catalog for the unified management of multimodal data, with robust support for integrating diverse external data sources

Unified Data Catalog

Built-in data catalog for the unified management of multimodal data, with robust support for integrating diverse external data sources

Unified Access Control

An RBAC-based unified permission model with a standardized access layer, delivering comprehensive access governance across the entire data lifecycle

Unified Access Control

An RBAC-based unified permission model with a standardized access layer, delivering comprehensive access governance across the entire data lifecycle

Intelligent Data Services

Built-in intelligent automation for small file compaction, snapshot cleanup, and overall data lifecycle management

Intelligent Data Services

Built-in intelligent automation for small file compaction, snapshot cleanup, and overall data lifecycle management

Open Ecosystem

Open integration with upper-layer compute engines, supporting Tencent Cloud services (e.g., EMR, DLC, TCHouse) and mainstream open-source ecosystems like Apache Spark and Flink

Open Ecosystem

Open integration with upper-layer compute engines, supporting Tencent Cloud services (e.g., EMR, DLC, TCHouse) and mainstream open-source ecosystems like Apache Spark and Flink

Multimodal Data Integration

Enables unified data lake ingestion and management by integrating structured data, unstructured AI data sources, and models

Batch-Stream Unified Table Format

A fully managed, batch-stream unified TCIceberg table format. It is fully compatible with Apache Iceberg while extending capabilities for streaming lakehouse scenarios

Unified Data Catalog

Built-in data catalog for the unified management of multimodal data, with robust support for integrating diverse external data sources

Unified Access Control

An RBAC-based unified permission model with a standardized access layer, delivering comprehensive access governance across the entire data lifecycle

Intelligent Data Services

Built-in intelligent automation for small file compaction, snapshot cleanup, and overall data lifecycle management

Open Ecosystem

Open integration with upper-layer compute engines, supporting Tencent Cloud services (e.g., EMR, DLC, TCHouse) and mainstream open-source ecosystems like Apache Spark and Flink

View All
Variety of solutions for your needs
Building a Lakehouse Architecture
Multimodal Data Lake Formation
Big Data & Machine Learning Integration
Building a Lakehouse Architecture
Build multi-scenario applications based on a unified data lake, such as Spark-based batch processing, Flink-based real-time pipelines, and TCHouse-based high-performance analytics. This resolves the fragmentation of traditional architectures that rely on isolated data silos for offline, real-time, and interactive analytics. Furthermore, by consolidating Lakehouse data assets via a unified metadata layer and providing intelligent optimization and acceleration services, it significantly enhances data maintenance and utilization efficiency for customers.
Key Features
  • Batch-Stream Unified Table Format (TCIceberg)
  • Unified Data Catalog
  • Multi-Engine Integration
Multimodal Data Lake Formation
Seamlessly integrate enterprise multimodal data—originally distributed across various heterogeneous systems—with TCLake's native data assets to achieve centralized management. This provides administrators with a globally visible asset governance console, while empowering upper-layer applications with standardized pan-domain data access, unified access control, and full-lifecycle data governance capabilities.
Key Features
  • Multimodal Data Management
  • Unified Access Control
  • External Data Source Integration
Big Data & Machine Learning Integration
Leveraging TCLake's multimodal data management capabilities and open engine ecosystem, customers can rapidly build integrated "Big Data + Machine Learning" applications. Training data preprocessed by upstream Big Data engines (like Apache Spark) can be registered directly into the unified metadata layer. Downstream AI training frameworks (such as PyTorch and TensorFlow) can then directly read this data. Once training is complete, the models can be registered back into TCLake for unified lifecycle management.
Key Features
  • Multimodal Data Catalog
  • Data+AI Multi-Engine Support
FAQS

Frequently

asked questions

What is the Multimodal Intelligent Data Lake (TCLake) service?

The Multimodal Intelligent Data Lake service (TCLake) is a next-generation open, intelligent, and converged AI data lake foundation launched by Tencent Cloud. It provides unified management covering both structured and unstructured data. With built-in features including a multimodal unified data catalog, a batch-stream unified table format, intelligent data management, and data acceleration services, it seamlessly integrates with Tencent Cloud services and mainstream open-source Data+AI ecosystem engines. This empowers enterprises to efficiently build a unified and cost-effective data lake infrastructure for the AI era.

What are the differences between the various data catalogs?

The metadata for Lakehouse, Volume, and Model catalogs is natively stored within the Tencent Cloud TCLake service. This includes metadata services compatible with Hive Metastore, as well as Volume-type data catalogs designed for the unified management of unstructured data. External data catalogs, on the other hand, establish connections with external data sources (such as MySQL) via JDBC or similar methods to retrieve metadata information from those sources in real time.

Which external data sources are currently supported?

Currently, Tencent Cloud DLC and TCHouse-D are supported. More data sources are actively being added.

How can I apply for the closed beta of the TCLake service?

The Multimodal Intelligent Data Lake (TCLake) service is currently in closed beta. This phase is open to invited users only, but you may fill out an application form to request access.

FAQS

Frequently

asked questions

What is the Multimodal Intelligent Data Lake (TCLake) service?

The Multimodal Intelligent Data Lake service (TCLake) is a next-generation open, intelligent, and converged AI data lake foundation launched by Tencent Cloud. It provides unified management covering both structured and unstructured data. With built-in features including a multimodal unified data catalog, a batch-stream unified table format, intelligent data management, and data acceleration services, it seamlessly integrates with Tencent Cloud services and mainstream open-source Data+AI ecosystem engines. This empowers enterprises to efficiently build a unified and cost-effective data lake infrastructure for the AI era.

What are the differences between the various data catalogs?

The metadata for Lakehouse, Volume, and Model catalogs is natively stored within the Tencent Cloud TCLake service. This includes metadata services compatible with Hive Metastore, as well as Volume-type data catalogs designed for the unified management of unstructured data. External data catalogs, on the other hand, establish connections with external data sources (such as MySQL) via JDBC or similar methods to retrieve metadata information from those sources in real time.

Which external data sources are currently supported?

Currently, Tencent Cloud DLC and TCHouse-D are supported. More data sources are actively being added.

How can I apply for the closed beta of the TCLake service?

The Multimodal Intelligent Data Lake (TCLake) service is currently in closed beta. This phase is open to invited users only, but you may fill out an application form to request access.