deallocate_blocks_cache
Applicability
Product |
Supported (√/x) |
|---|---|
Atlas 350 Accelerator Card |
x |
√ |
|
√ |
|
x |
|
x |
|
x |
Note: For
Function Description
In the PagedAttention scenario, releases the cache allocated by allocate_blocks_cache.
Prototype
1 | deallocate_blocks_cache(cache: Cache) |
Parameters
Parameter |
Data Type |
Description |
|---|---|---|
cache |
Cache to be released. |
Example
1 2 3 | from llm_datadist import BlocksCacheKey ... cache_manager.deallocate_blocks_cache(blocks_cache) |
Returns
In normal cases, no value is returned.
If the input data type is incorrect, the TypeError or ValueError exception is reported.
If the execution time exceeds the value of sync_kv_timeout, an LLMException is thrown.
Restrictions
None