intellicrack.providers.model_loader

Model loading utilities with quantization and caching for local transformers.

This module provides model loading, caching, and memory management for HuggingFace Transformers models optimized for Intel XPU and CPU inference.

validate_local_checkpoint(model_id)[source]

Reject a local checkpoint whose sharded-weight index escapes its folder.

transformers resolves the shard file names in a sharded checkpoint’s *.index.json by joining each weight_map value directly onto the checkpoint folder, so a hostile checkpoint can point a shard at ../ paths, absolute paths, or a named pipe and have the loader read arbitrary files or block indefinitely (CVE-2026-69112, unreachable in accelerate here but reachable through the transformers local-folder loader). This runs before any from_pretrained call and validates every shard entry of every index found in a local checkpoint directory.

Hugging Face Hub repository ids (anything that is not an existing local directory) are left untouched: the SDK resolves and caches those itself.

Parameters:

model_id (str) – The configured model identifier or local checkpoint path.

Return type:

None

class LoadedModel[source]

Bases: object

A loaded model with its tokenizer and metadata.

model: PreTrainedModel
tokenizer: PreTrainedTokenizerBase
device: torch.device
dtype: str
memory_usage_bytes: int
model_id: str
load_time_seconds: float
__init__(model, tokenizer, device, dtype, memory_usage_bytes, model_id, load_time_seconds)
Parameters:
  • model (PreTrainedModel)

  • tokenizer (PreTrainedTokenizerBase)

  • device (torch.device)

  • dtype (str)

  • memory_usage_bytes (int)

  • model_id (str)

  • load_time_seconds (float)

Return type:

None

class ModelConfig[source]

Bases: object

Configuration for model loading.

Variables:
  • model_id (str) – HuggingFace model identifier or local path.

  • dtype (Literal['auto', 'float32', 'float16', 'bfloat16', 'int8', 'int4']) – Data type for the model.

  • device (Literal['xpu', 'cpu', 'auto']) – Target device.

  • max_memory_bytes (int) – Maximum memory to use.

  • trust_remote_code (bool) – Whether to trust remote code.

  • use_flash_attention (bool) – Whether to use flash attention if available.

  • quantization_config (dict[str, object] | None) – Optional quantization configuration.

  • revision (str | None) – Git revision (commit hash, tag, or branch) to pin downloads to a specific snapshot of the model repository. When None, HuggingFace defaults to the main branch.

model_id: str
dtype: Literal['auto', 'float32', 'float16', 'bfloat16', 'int8', 'int4'] = 'auto'
device: Literal['xpu', 'cpu', 'auto'] = 'auto'
max_memory_bytes: int = 12884901888
trust_remote_code: bool = False
use_flash_attention: bool = False
quantization_config: dict[str, object] | None = None
revision: str | None = None
__init__(model_id, dtype='auto', device='auto', max_memory_bytes=12884901888, trust_remote_code=False, use_flash_attention=False, quantization_config=None, revision=None)
Parameters:
  • model_id (str)

  • dtype (Literal['auto', 'float32', 'float16', 'bfloat16', 'int8', 'int4'])

  • device (Literal['xpu', 'cpu', 'auto'])

  • max_memory_bytes (int)

  • trust_remote_code (bool)

  • use_flash_attention (bool)

  • quantization_config (dict[str, object] | None)

  • revision (str | None)

Return type:

None

class ModelCache[source]

Bases: object

LRU cache for loaded models with memory limit enforcement.

Maintains an LRU cache of loaded models, automatically evicting least recently used models when the memory limit is exceeded.

__init__(max_memory_bytes=10737418240)[source]

Initialize the ModelCache with a memory limit.

Parameters:

max_memory_bytes (int) – Maximum memory in bytes allowed for cached models.

Return type:

None

property max_memory_bytes: int

The maximum memory limit.

Returns:

Maximum memory in bytes allowed for cached models.

Return type:

int

get(model_id, dtype, device_type)[source]

Get a model from cache.

Parameters:
  • model_id (str) – The model identifier.

  • dtype (str) – The data type.

  • device_type (str) – The device type.

Returns:

The cached LoadedModel or None if not cached.

Return type:

LoadedModel | None

put(loaded_model)[source]

Put a model into cache.

Parameters:

loaded_model (LoadedModel) – The loaded model to cache.

Return type:

None

remove(model_id, dtype, device_type)[source]

Remove a model from cache.

Parameters:
  • model_id (str) – The model identifier.

  • dtype (str) – The data type.

  • device_type (str) – The device type.

Returns:

True if model was removed, False if not found.

Return type:

bool

clear()[source]

Clear all cached models.

Return type:

None

get_memory_usage()[source]

Get current memory usage.

Returns:

Current memory usage in bytes.

Return type:

int

estimate_model_memory(model_id, dtype='float16', *, include_activations=True)[source]

Estimate memory required for a model.

Parameters:
  • model_id (str) – HuggingFace model identifier or path.

  • dtype (Literal['auto', 'float32', 'float16', 'bfloat16', 'int8', 'int4']) – Data type for the model.

  • include_activations (bool) – Include activation memory overhead.

Returns:

Estimated memory in bytes.

Return type:

int

select_dtype_for_memory(model_id, available_memory_bytes, preferred_dtype='auto')[source]

Select appropriate dtype to fit model in available memory.

Parameters:
  • model_id (str) – HuggingFace model identifier.

  • available_memory_bytes (int) – Available memory in bytes.

  • preferred_dtype (Literal['auto', 'float32', 'float16', 'bfloat16', 'int8', 'int4']) – Preferred dtype if it fits.

Returns:

Selected dtype that should fit in memory.

Return type:

DtypeOption

load_model_for_xpu(config, cache=None)[source]

Load a model optimized for Intel XPU.

Parameters:
Returns:

LoadedModel with model, tokenizer, and metadata.

Return type:

LoadedModel

Raises:
load_model_for_cpu(config, cache=None)[source]

Load a model for CPU inference.

Parameters:
Returns:

LoadedModel with model, tokenizer, and metadata.

Return type:

LoadedModel

Raises:
get_global_model_cache()[source]

Get the global model cache singleton.

Returns:

The global ModelCache instance.

Return type:

ModelCache

set_global_cache_size(max_memory_bytes)[source]

Set the global cache size limit.

Parameters:

max_memory_bytes (int) – Maximum memory for the cache.

Return type:

None

clear_global_cache()[source]

Clear the global model cache.

Return type:

None