intellicrack.providers.grok
X.AI Grok API provider implementation.
This module provides integration with X.AI’s Grok models for chat completion and tool/function calling. Grok uses an OpenAI-compatible API, so this implementation leverages the OpenAI SDK with a custom base URL.
- class GrokMessage[source]
Bases:
TypedDictGrok message structure.
- content: str | list[GrokMessageContent] | None
- class GrokProvider[source]
Bases:
LLMProviderBaseX.AI Grok API provider implementation.
Provides integration with X.AI’s Grok models including support for tool/function calling and streaming responses. Uses the OpenAI SDK with a custom base URL for API compatibility.
- Variables:
- property dialect: ApiDialect
The wire format this provider speaks.
- Returns:
Always
ApiDialect.CHAT_COMPLETIONS.- Return type:
- async connect(credentials)[source]
Connect to X.AI Grok API.
- Parameters:
credentials (ProviderCredentials) – Must contain api_key. Optionally api_base for custom URL, and
timeoutto replace the SDK’s default request timeout.- Raises:
AuthenticationError – If API key is invalid.
ProviderError – If connection fails.
- Return type:
None
- async list_models()[source]
Dynamically fetch available models from Grok.
- Returns:
List of available Grok models.
- Return type:
- Raises:
ProviderError – If not connected.
- async chat(messages, model, tools=None, temperature=0.7, max_tokens=4096, tool_choice=None, thinking=None, *, enable_cache=False)[source]
Send a chat completion request to Grok.
thinkingis honoured on Grok models that expose it through the OpenAI-compatiblereasoning_effortparameter (grok-4-multi-agent). Grok-4 / Grok-4-fast reason automatically, so the parameter is intentionally omitted there.- Parameters:
model (str) – Model ID to use.
tools (list[ToolDefinition] | None) – Available tools for function calling.
temperature (float) – Sampling temperature.
max_tokens (int) – Maximum tokens in response.
tool_choice (ToolChoice | None) – How the model should select tools.
thinking (ThinkingConfig | None) – Extended thinking configuration. Forwarded as
reasoning_effortforgrok-*-multi-agentmodels.enable_cache (bool) – Whether to enable prompt caching. Grok caches automatically when the same prompt prefix is reused; the parameter is logged for symmetry.
- Returns:
Tuple of (assistant message, tool calls if any).
- Return type:
- Raises:
ProviderError – If not connected or request fails.
- async chat_stream(messages, model, tools=None, temperature=0.7, max_tokens=4096, tool_choice=None, thinking=None, *, enable_cache=False)[source]
Stream a chat completion response from Grok.
- Parameters:
model (str) – Model ID to use.
tools (list[ToolDefinition] | None) – Available tools for function calling.
temperature (float) – Sampling temperature.
max_tokens (int) – Maximum tokens in response.
tool_choice (ToolChoice | None) – How the model should select tools.
thinking (ThinkingConfig | None) – Extended thinking configuration. Forwarded as
reasoning_effortforgrok-*-multi-agentmodels; ignored on grok-4 / grok-4-fast (auto-reasoning).enable_cache (bool) – Whether to enable prompt caching. Grok caches automatically; the parameter is logged for symmetry.
- Yields:
str – Text chunks as they arrive.
- Raises:
ProviderError – If not connected or the request fails. An invalid key raises the
AuthenticationErrorsubclass and a transient rate limit theRateLimitErrorsubclass.- Return type:
AsyncIterator[str]