deallocate_cache

Applicable Products

Product

Supported (Yes/No)

Ascend 950PR/Ascend 950DT

No

Atlas A3 training products/Atlas A3 inference products

Yes

Atlas A2 training products/Atlas A2 inference products

Yes

Atlas 200I/500 A2 inference products

No

Atlas inference products

No

Atlas training products

No

Note: For the Atlas A2 training products/Atlas A2 inference products, only the Atlas 800I A2 inference server and A200I A2 Box heterogeneous subrack are supported.

Function Description

Deallocates a cache.

If the cache is associated with a cache key during allocation, the actual release is delayed until all cache keys are pulled or remove_cache_key is executed.

Prototype

1
deallocate_cache(cache: KvCache)

Parameters

Parameter

Data Type

Value Description

cache

KvCache

KV cache to be deallocated.

Example

1
kv_cache_manager.deallocate_cache(kv_cache)

Returns

In normal cases, no value is returned.

If a parameter is incorrect, a TypeError or ValueError may be thrown.

If the execution time exceeds the value of sync_kv_timeout, an LLMException is thrown.

Constraints

  • If KvCache does not exist or has been released, this operation is a no-operation.
  • This API does not support concurrent calls.