intellicrack.providers.model_loader
Model loading utilities with quantization and caching for local transformers.
This module provides model loading, caching, and memory management for HuggingFace Transformers models optimized for Intel XPU and CPU inference.
- validate_local_checkpoint(model_id)[source]
Reject a local checkpoint whose sharded-weight index escapes its folder.
transformersresolves the shard file names in a sharded checkpoint’s*.index.jsonby joining eachweight_mapvalue directly onto the checkpoint folder, so a hostile checkpoint can point a shard at../paths, absolute paths, or a named pipe and have the loader read arbitrary files or block indefinitely (CVE-2026-69112, unreachable inacceleratehere but reachable through the transformers local-folder loader). This runs before anyfrom_pretrainedcall and validates every shard entry of every index found in a local checkpoint directory.Hugging Face Hub repository ids (anything that is not an existing local directory) are left untouched: the SDK resolves and caches those itself.
- Parameters:
model_id (str) – The configured model identifier or local checkpoint path.
- Return type:
None
- class LoadedModel[source]
Bases:
objectA loaded model with its tokenizer and metadata.
- model: PreTrainedModel
- tokenizer: PreTrainedTokenizerBase
- device: torch.device
- class ModelConfig[source]
Bases:
objectConfiguration for model loading.
- Variables:
model_id (str) – HuggingFace model identifier or local path.
dtype (Literal['auto', 'float32', 'float16', 'bfloat16', 'int8', 'int4']) – Data type for the model.
device (Literal['xpu', 'cpu', 'auto']) – Target device.
max_memory_bytes (int) – Maximum memory to use.
trust_remote_code (bool) – Whether to trust remote code.
use_flash_attention (bool) – Whether to use flash attention if available.
quantization_config (dict[str, object] | None) – Optional quantization configuration.
revision (str | None) – Git revision (commit hash, tag, or branch) to pin downloads to a specific snapshot of the model repository. When
None, HuggingFace defaults to themainbranch.
- __init__(model_id, dtype='auto', device='auto', max_memory_bytes=12884901888, trust_remote_code=False, use_flash_attention=False, quantization_config=None, revision=None)
- Parameters:
- Return type:
None
- class ModelCache[source]
Bases:
objectLRU cache for loaded models with memory limit enforcement.
Maintains an LRU cache of loaded models, automatically evicting least recently used models when the memory limit is exceeded.
- __init__(max_memory_bytes=10737418240)[source]
Initialize the ModelCache with a memory limit.
- Parameters:
max_memory_bytes (int) – Maximum memory in bytes allowed for cached models.
- Return type:
None
- property max_memory_bytes: int
The maximum memory limit.
- Returns:
Maximum memory in bytes allowed for cached models.
- Return type:
- get(model_id, dtype, device_type)[source]
Get a model from cache.
- Parameters:
- Returns:
The cached LoadedModel or None if not cached.
- Return type:
LoadedModel | None
- put(loaded_model)[source]
Put a model into cache.
- Parameters:
loaded_model (LoadedModel) – The loaded model to cache.
- Return type:
None
- estimate_model_memory(model_id, dtype='float16', *, include_activations=True)[source]
Estimate memory required for a model.
- select_dtype_for_memory(model_id, available_memory_bytes, preferred_dtype='auto')[source]
Select appropriate dtype to fit model in available memory.
- Parameters:
- Returns:
Selected dtype that should fit in memory.
- Return type:
DtypeOption
- load_model_for_xpu(config, cache=None)[source]
Load a model optimized for Intel XPU.
- Parameters:
config (ModelConfig) – Model configuration.
cache (ModelCache | None) – Optional model cache.
- Returns:
LoadedModel with model, tokenizer, and metadata.
- Return type:
- Raises:
RuntimeError – If model loading fails.
ImportError – If required packages are not installed.
- load_model_for_cpu(config, cache=None)[source]
Load a model for CPU inference.
- Parameters:
config (ModelConfig) – Model configuration.
cache (ModelCache | None) – Optional model cache.
- Returns:
LoadedModel with model, tokenizer, and metadata.
- Return type:
- Raises:
RuntimeError – If model loading fails.
ImportError – If required packages are not installed.
- get_global_model_cache()[source]
Get the global model cache singleton.
- Returns:
The global ModelCache instance.
- Return type: