Big Data Intelligent Agent Workbench DataBuddy is an Agent-Native, fully managed Data + AI integrated data intelligence platform introduced by Tencent Cloud. DataBuddy combines data computing and AI Intelligent Agent capacity. Through unified metadata and semantic layer, it provides enterprises with end-to-end full-link capacity from data access, data engineering, data science, data analysis to data governance, evolving big data platforms from "humans operate tools" to "AI works, humans gatekeep".
All Tencent Cloud DataBuddy APIs introduced in this chapter comply with the latest OpenAPI 3.0 specification.
You can call APIs to operate Tencent Cloud DataBuddy, such as creating and managing workspaces, creating and resizing computational resources, managing data source connections, submitting and querying SQL jobs, creating and running workflows, querying job running instances and running states, managing assets like Catalog/Schema/Table/Volume in the data directory, registering and deploying model services, and querying data quality and governance results, so as to integrate DataBuddy capabilities into your own scheduling system, ops platform, or business applications.
Common terminology for Tencent Cloud DataBuddy API interfaces: see the table below
| Term | Description |
|---|---|
| Workspace | The top-level resource and permission isolation container of DataBuddy. It hosts development files, tasks, computational resources, members, and configurations, usually divided by team, project, or environment. The workspace needs to be specified when calling most service interfaces. |
| Region | The physical location of the data center, which determines the physical location of data storage and compute resources. The region cannot be modified once a workspace is created. Specify it through the common parameter Region when calling APIs. |
| Data Catalog | The top-level container in the level 3 namespace (Catalog.Schema.Table). It is used for unified management of metadata, permissions, and data assets across workspaces, serving as the core boundary for data governance and access control. |
| Schema | A second-level namespace in a data catalog. It is used for logical grouping and isolation of tables, views, functions, and other objects. It usually corresponds to a business domain, project, or environment. |
| Table | The basic storage unit for structured data. It includes column definitions and data types, and supports metadata query and management through the API. |
| Volume | A storage unit oriented toward unstructured and semi-structured data (such as images, audio and video, PDF, JSON, Parquet files). Located under Schema, included in unified permission management, and directly accessible by Notebook, jobs, and model training. |
| Computational resource | Cluster resources in DataBuddy that handle computational load, divided into job cluster, interactive cluster, platform cluster, and real-time cluster. Used to run tasks such as SQL, Notebook, Python, Ray, and data access. |
| CU (Compute Unit) | A unified measurement unit for computational resources. 1 CU ≈ 1 core 4 GB memory. Used for specification quota settings and metered billing of computational resources. |
| Workflow | A unit for orchestrating and scheduling data tasks. It connects nodes such as Notebook, SQL, data access, and model training in DAG format, supporting dependency management, scheduled triggering, and alerts. |
| Task instance | A run history generated only once when a task in the workflow is triggered by a scheduling interval or manually. Used for querying the running state, logs, and execution result. |
| Backfill | The ability to re-execute tasks against a historical time range. Used for fixing data missing, correcting logic errors, or initializing historical partitions. Supports batch backfill by Business Date. |
| Data source | Connection configuration for external systems (database, data warehouse, message queue, object storage, API, etc.). Used for offline integration, real-time integration, and federated query. It can be used only after passing the connectivity test. |
| Realtime Ingestion | A low-latency data synchronization method based on technologies such as CDC (change data capture). It synchronizes data changes from the source system to the data lake in real time, supporting real-time data warehouse and real-time analytics scenarios. |
| Offline integration (Batch Ingestion) | Synchronize external data sources to the data lake in batches on a cycle. Suitable for large-scale data integration with low timeliness requirements. |
| Semantic Model | A business semantic abstraction layer located above the physical data layer. It provides a unified definition of metrics, dimensions, entities and their relationships, ensuring consistent semantic caliber for BI analysis, natural language querying and AI applications. |
| Metric | A measurable value definition oriented toward business scenarios (for example, GMV, retention rate, number of active users). It includes caliber, computation logic, dimensions, and aggregation methods for unified calls by BI, Agent, and downstream applications. |
| Model | A machine learning model asset that includes model files, dependencies, signatures, and metadata. It supports version management, lineage tracing, and lifecycle management, and can be registered to a model repository for reasoning or retraining. |
| Model Service | The ability to deploy registered models as online or batch inference services. Provides REST APIs, auto scaling, traffic routing, and monitoring for integrating model capabilities into business systems. |
| Intelligent Agent | An AI app unit that can plan, call tools, and has memory capacity. It can connect to metrics, data, models, and external systems to complete complex business tasks and automated processes. |
| Buddy | A built-in AI intelligent assistant in the platform, covering data development, analysis, governance, O&M, and other scenarios. It can assist users in writing SQL, explaining code, troubleshooting tasks, and generating documents through natural language. |
| OBO (On Behalf Of) | A permission model where an Agent executes On Behalf Of a user's real identity. Both the Agent and API calls reuse the caller's privileges On the data platform to prevent unauthorized access. |
| Governed Tags | Controlled Tags applied to objects such as catalogs, schemas, tables, and columns in the data Catalog. Used for data categorization, sensitive data identification, access policies, and compliance auditing. |
| RequestId | Unique request ID returned for each API call. Used for problem localization and ticket troubleshooting. We recommend that you keep a complete record of it in business logs. |
When calling the Tencent Cloud DataBuddy API, please note the following limits:
Region, and only supports regions where the current account has activated the DataBuddy service. Cross-region resources cannot be operated across each other, and the region of a workspace cannot be modified after creation.Offset and Limit, and the number of entries returned at once has an upper limit. Use loop pagination to obtain the complete data. Do not depend on a single full return.You can use the API Explorer tool to call APIs online.
This document uses creating a workflow and viewing workflow info as an example. The steps to make an API call through the API Explorer Tool are as follows:
WorkspaceId of the target workspace. Subsequently, all workflow APIs need to input this parameter.WorkspaceId and BaseInfo (where WorkflowName is required and unique within the workspace). For scheduled dispatch, configure a Cron expression via Trigger. The API response returns the WorkflowId of the new workflow.WorkflowNameKeyword for name fuzzy search and paginate through PageNumber and PageSize.WorkspaceId and WorkflowId to view the complete definition of the workflow (basic information, task nodes, scheduling, alarms, parameters, tags, etc.).Create workflow:
{
"WorkspaceId": "ws-xxxxxxxx",
"BaseInfo": {
"WorkflowName": "daily_etl_pipeline",
"Description": "Daily ETL main process"
},
"Trigger": [
{
"Type": "CRON",
"CronExpression": "0 0 2 * * ?"
}
]
}
Output result:
{
"Response": {
"Data": {
"WorkflowId": "wf-xxxxxxxx"
},
"RequestId": "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
}
}
Workflow lifecycle APIs are as follows:
| Scenario | API |
|---|---|
| Create / update / delete workflow | CreateWorkflow, UpdateWorkflow, DeleteWorkflow |
| Query workflow list/details | ListWorkflows, GetWorkflow |
| Run / Rerun / Terminate | RunWorkflow, RerunWorkflowRun, KillWorkflowRun |
| Query the run record list / details | ListWorkflowRuns, GetWorkflowRun |
| Query task run list / details | ListWorkflowTaskRuns, GetWorkflowTaskRun |
Apakah halaman ini membantu?
Anda juga dapat Menghubungi Penjualan atau Mengirimkan Tiket untuk meminta bantuan.
masukan