tencent cloud

Cost management

Download
Focus Mode
Font Size
Last updated: 2026-09-30 18:41:47
AI-Translated & Reviewed

Overview

Optimize costs through multi-scenario model selection and an asynchronous compression policy, achieving faster response times and lower usage costs while maintaining performance. CodeBuddy Code consumes tokens on every interaction. Costs vary depending on codebase size, query complexity, and conversation length. This document describes how to track costs, the multi-scenario model mechanism, and how to reduce Token consumption.

Tracking Costs

Using the /cost Command

The /cost command provides detailed Token usage statistics for the current session:
/cost
⎿ Total duration (API): 9m 35.6s
Total duration (wall): 22m 14.9s
Total code changes: 0 lines added, 0 lines removed
Usage by model:
claude-sonnet-4: 875.5k input, 11.7k output, 714.3k cache read, 0 cache write

Using the /context Command

The /context command analyzes the current context usage and shows the size distribution of different context types:
> /context
⎿ Context Usage
⛁ ⛁ ⛁ ⛁ ⛁ ⛁ ⛁ ⛁ ⛁ ⛁ glm-4.7 · 38.1k/200.0k tokens (19.1%)
⛁ ⛁ ⛁ ⛁ ⛁ ⛁ ⛁ ⛁ ⛁ ⛁
⛁ ⛶ ⛶ ⛶ ⛶ ⛶ ⛶ ⛶ ⛶ ⛶ ⛁ System prompt: 2.1k tokens (1.1%)
⛶ ⛶ ⛶ ⛶ ⛶ ⛶ ⛶ ⛶ ⛶ ⛶ ⛁ System tools: 16.4k tokens (8.2%)
⛶ ⛶ ⛶ ⛶ ⛶ ⛶ ⛶ ⛶ ⛶ ⛶ ⛁ Memory files: 3.7k tokens (1.9%)
⛶ ⛶ ⛶ ⛶ ⛶ ⛶ ⛶ ⛶ ⛶ ⛶ ⛁ Messages: 15.9k tokens (7.9%)
⛶ ⛶ ⛶ ⛶ ⛶ ⛶ ⛶ ⛶ ⛶ ⛶ ⛶ Free space: 145.9k (72.9%)
⛶ ⛶ ⛶ ⛶ ⛶ ⛶ ⛶ ⛶ ⛶ ⛶ ⛝ Autocompact buffer: 16.0k tokens (8.0%)
⛶ ⛶ ⛶ ⛶ ⛶ ⛶ ⛶ ⛶ ⛶ ⛶
⛶ ⛶ ⛶ ⛶ ⛝ ⛝ ⛝ ⛝ ⛝ ⛝

Memory files · /memory
└ /Users/yangsubo/.codebuddy/CODEBUDDY.md (User): 18 tokens
└ /Users/yangsubo/CODEBUDDY.md (Project): 15 tokens
└ /Users/yangsubo/workspace/genie/CODEBUDDY.md (Project): 1.2k tokens
└ /Users/yangsubo/workspace/genie/packages/agent-cli/CODEBUDDY.md (Project): 2.5k tokens

Skills and slash commands · /skills

Project
└ release: 1.1k tokens
└ gen-drawio: 846 tokens
└ task-manager: 815 tokens
└ task-add: 730 tokens
└ task-done: 525 tokens
└ mr: 519 tokens
└ task-start: 493 tokens
└ my-task: 238 tokens
└ task-list: 233 tokens
└ security-review: 30 tokens
You can use /context to quickly identify which content is taking up a large amount of context space, allowing for targeted optimization.

Multi-Scenario Model Mechanism

Different task scenarios have different requirements for model capabilities. Simple tasks such as file search and quick queries can be completed with lightweight models, while complex tasks such as architecture design and multi-step reasoning require more powerful reasoning models.
By automatically selecting different models for different scenarios, you can achieve the following:
Optimal performance: Use high-capability models for complex tasks to ensure quality.
Faster speed: Simple tasks use lightweight models for quicker responses.
Lower cost: Avoid using expensive high-end models for simple tasks.

Scenario Type

Scenario Type
Description
Typical Use Case
default
Default model, balancing performance and cost
General programming tasks and code writing
lite
Lightweight and fast model with low cost and high speed
File search, simple queries, and quick operations
reasoning
Reasoning-enhanced model with powerful analysis capabilities
Complex analysis, architectural decisions, and multi-step reasoning

Automatic Model Selection

CodeBuddy Code resolves the appropriate scenario model based on the task type. For example, the built-in declaration for Explore is lite, allowing code search to be handled by a faster and more cost-effective model. For planning tasks, reasoning can be used.
Use /model:lite and /model:reasoning to map lite and reasoning to specific models. To configure a built-in subagent individually, select a specific model or scenario variant through /agents. Project settings only override subagents or scenario keys with the same name and do not delete other global mappings.
Agent tools support specifying the scenario type through the model parameter:
default: Continue using the default orchestration composed of per-subagent settings, built-in declarations, and the main model.
lite: Fast and low-cost, suitable for simple searches and quick file operations.
reasoning: Enhances reasoning capabilities, suitable for complex analysis and architectural decisions.

Reducing Token Consumption

Token costs grow with context size: the larger the context processed by CodeBuddy Code, the more tokens are consumed. CodeBuddy Code automatically optimizes costs through Prompt caching, which reduces the cost of repetitive content such as system prompts, and automatic compaction, which compresses conversation history when the context limit is approached.
The following strategies help you maintain a smaller context and reduce the cost per message.

Proactively Managing Context

Use /cost to check the current Token usage.
Cleanup between tasks: Use /clear to start fresh when switching to unrelated work. Stale context wastes tokens in every subsequent message. Use /rename before clearing so you can return later via /resume.
Add custom compaction instructions: /compact Focus on code samples and API usage tells CodeBuddy Code what to retain during compaction.
You can also customize the compaction behavior in CODEBUDDY.md:
# Compact instructions

When you are using compact, please focus on test output and code changes.

Asynchronous Compression Policy

When the conversation history approaches the context limit, the system automatically compacts it:
Automatic triggering: Automatically starts compaction when the context approaches the limit.
Background execution: The compaction process runs asynchronously in the background without blocking user operations.
Smart summary: Retains key information and compresses redundant content.
Seamless transition: Users enjoy an "infinite context" experience without noticing it.
The following key information is retained during compaction: code change records, important decision points, explicit user preferences and instructions, and key context of the current task.

Selecting an Appropriate Model

Select a model based on task complexity. Use /model to switch the main model, or use /model:lite / /model:reasoning to configure lite / reasoning. Use /agents to select a scenario variant or a specific model for a particular built-in subagent.
Use lite for simple tasks: file search, quick queries, and code formatting.
Use reasoning for complex tasks: architectural design, performance optimization, and complex debugging.
Use default for general tasks: daily coding and feature implementation.

Reducing MCP Server Overhead

Each MCP server adds its tool definitions to the context even when idle. Run /mcp to view configured servers.
Prefer CLI tools: Tools such as gh, aws, and gcloud consume less context than MCP servers because they do not add persistent tool definitions. CodeBuddy Code can run CLI commands directly without additional overhead.
Disable unused servers: Run /mcp to view and disable unused servers.

Delegating Detailed Operations to Subagents

Running tests, obtaining documentation, or processing log files can consume a large amount of context. Delegate these operations to subagents so that detailed output remains in the subagent's context and only a summary returns to the main conversation.

Writing Precise Prompts

Vague requests such as "improve this codebase" trigger broad scanning. Precise requests such as "add input validation to the login function in auth.ts" enable CodeBuddy Code to work efficiently with minimal file reads.

Efficient Workflows for Complex Tasks

For longer or more complex tasks, the following habits help avoid wasting tokens by heading in the wrong direction:
Use plan mode for complex tasks: Press Shift+Tab (Alt+M is also supported on Windows) to enter plan mode. CodeBuddy Code explores the codebase and proposes a plan for your approval, avoiding costly rework if the initial direction is wrong.
Correct course early: If CodeBuddy Code starts heading in the wrong direction, press Escape to stop immediately. Use /rewind or double-press Escape to restore the conversation and code to a previous checkpoint.
Provide verification targets: Include test cases, screenshots, or expected outputs in your prompts. When CodeBuddy Code can self-verify its work, it can identify issues before you need to request fixes.
Test incrementally: Write one file, test it, and then continue. This helps identify issues early while they are still easy to fix.

Background Token Consumption

CodeBuddy Code also consumes tokens for certain background features even when idle:
Conversation summarization: A background task summarizes previous conversations for the --resume feature.
Prompt prediction: Predicts the most likely next Prompt input based on historical conversation information.
These background processes consume a small number of tokens even without active interaction.

References

Subagents - Use subagents to isolate high-consumption operations.
MCP - Manage MCP server overhead.
Model - Learn about available model options.


Help and Support

Was this page helpful?

Help us improve! Rate your documentation experience in 5 mins.

Feedback