deallocate_blocks_cache

Applicability

Product

Supported (√/x)

Atlas 350 Accelerator Card

x

Atlas A3 training product/Atlas A3 inference product

Atlas A2 training product/Atlas A2 inference product

Atlas 200I/500 A2 inference product

x

Atlas inference product

x

Atlas training product

x

Note: For Atlas A2 training product/Atlas A2 inference product, only the Atlas 800I A2 inference server and A200I A2 Box heterogeneous subrack are supported.

Function Description

In the PagedAttention scenario, releases the cache allocated by allocate_blocks_cache.

Prototype

1
deallocate_blocks_cache(cache: Cache)

Parameters

Parameter

Data Type

Description

cache

Cache

Cache to be released.

Example

1
2
3
from llm_datadist import BlocksCacheKey
...
cache_manager.deallocate_blocks_cache(blocks_cache)

Returns

In normal cases, no value is returned.

If the input data type is incorrect, the TypeError or ValueError exception is reported.

If the execution time exceeds the value of sync_kv_timeout, an LLMException is thrown.

Restrictions

None