get_cache_tensors

Applicable Products

Product

Supported (Yes/No)

Ascend 950PR/Ascend 950DT

No

Atlas A3 training products/Atlas A3 inference products

Yes

Atlas A2 training products/Atlas A2 inference products

Yes

Atlas 200I/500 A2 inference products

No

Atlas inference products

No

Atlas training products

No

Note: For the Atlas A2 training products/Atlas A2 inference products, only the Atlas 800I A2 inference server and A200I A2 Box heterogeneous subrack are supported.

Function Description

Obtains a cache tensor.

Prototype

1
get_cache_tensors(cache: KvCache, tensor_index: int = 0) -> List[Tensor]

Parameters

Parameter

Data Type

Value Description

cache

KvCache

KvCache where the tensor to be obtained is located.

tensor_index

int

Index of the tensor to be obtained in the cache.

Example

1
tensors = kv_cache_manager.get_cache_tensors(kv_cache, 0)

Returns

In normal cases, List[Tensor] is returned.

If a parameter is incorrect, a TypeError or ValueError may be thrown.

If the execution time exceeds the value of sync_kv_timeout, an LLMException is thrown.

Constraints

This API does not support concurrent calls.