PullKvCache
Applicability
Product |
Supported (Yes/No) |
|---|---|
Atlas 350 Accelerator Card |
No |
Yes |
|
Yes |
|
No |
|
No |
|
No |
Note: For the
Description
Pulls the cache from the remote node to the local cache. This API can be called only when the role is Decoder.
Prototype
1 2 3 4 5 | Status PullKvCache(const CacheIndex &src_cache_index, const Cache &dst_cache, uint32_t batch_index = 0U, int64_t size = -1, const KvCacheExtParam &ext_param = {}) |
Parameters
Parameter |
Input/Output |
Description |
|---|---|---|
src_cache_index |
Input |
Index of the remote source cache. |
dst_cache |
Input |
Local destination cache. |
batch_index |
Input |
Index of the local destination batch. |
size |
Input |
If this parameter is set to an integer greater than 0, it indicates the size of the data to be pulled. If this parameter is set to -1, it indicates that the entire data is pulled. The default value is -1. |
ext_param |
Input |
The difference between second and first in src_layer_range must equal that in dst_layer_range. The default values of first and second in both src_layer_range and dst_layer_range are -1, indicating all layers. The valid range for each index is [0, maximum available layer index], and first must be less than or equal to second. The maximum available layer index is calculated as follows: (CacheDesc::num_tensors / KvCacheExtParam::tensor_num_per_layer) - 1 tensor_num_per_layer can range from 1 to the total number of tensors in the cache, with a default value of 2. When src_layer_range or dst_layer_range uses non-default values, tensor_num_per_layer can either remain at its default value or be set to another value, which must be divisible by the total number of tensors in the cache. |
Example
1 2 3 4 5 | CacheIndex cache_index; cache_index.cluster_id = 0; cache_index.cache_id = cached_tensors.cache_id; cache_index.batch_index = 0; Status ret = llm_datadist.PullKvCache(cache_index, cache) |
Returns
- LLM_SUCCESS: Success.
- LLM_PARAM_INVALID: Incorrect parameter.
- LLM_NOT_YET_LINK: No link was established with the remote cluster.
- LLM_TIMEOUT: Pull timed out.
- LLM_KV_CACHE_NOT_EXIST: The local or remote KV cache does not exist.
- Other values: Failure.
Constraints
Before calling this API, call the Initialize API to complete initialization. dst_cache must be the cache allocated via the AllocateCache API.