intellicrack.providers.openrouter

OpenRouter API provider implementation.

This module provides integration with OpenRouter which provides access to many different LLM providers through a unified API.

class OpenRouterProvider[source]

Bases: LLMProviderBase

OpenRouter API provider implementation.

Provides access to many different LLM models through OpenRouter’s unified API interface.

Variables:
  • BASE_URL (str) – OpenRouter unified LLM API base URL.

  • TOOL_COUNT_CAP (int | None) – Maximum number of flattened tool functions sent in a single function-calling request, matching the tightest common limit among OpenAI-compatible backends OpenRouter routes to (e.g. gpt-4o-mini).

BASE_URL: str = 'https://openrouter.ai/api/v1'
TOOL_COUNT_CAP: int | None = 128
__init__()[source]

Initialize the OpenRouterProvider instance.

Return type:

None

property name: str

The provider instance id.

Returns:

The openrouter built-in provider id.

Return type:

str

property dialect: ApiDialect

The wire format this provider speaks.

Returns:

Always ApiDialect.CHAT_COMPLETIONS.

Return type:

ApiDialect

async connect(credentials)[source]

Connect to OpenRouter API.

HTTP errors raised by the connect probe are translated to Intellicrack typed errors by LLMProviderBase._raise_typed_for_status().

Parameters:

credentials (ProviderCredentials) – Must contain api_key. api_base optionally overrides the API endpoint for proxy or self-hosted OpenRouter-compatible gateways; it defaults to BASE_URL.

Raises:
Return type:

None

async disconnect()[source]

Disconnect from OpenRouter API.

Return type:

None

async list_models()[source]

Dynamically fetch available models from OpenRouter.

Returns:

List of available models.

Return type:

list[ModelInfo]

Raises:

ProviderError – If not connected.

async chat(messages, model, tools=None, temperature=0.7, max_tokens=4096, tool_choice=None, thinking=None, *, enable_cache=False)[source]

Send a chat completion request through OpenRouter.

HTTP errors are translated to Intellicrack typed errors by LLMProviderBase._raise_typed_for_status(). Transient rate-limit failures are retried via LLMProviderBase._retry_with_backoff(). enable_cache attaches OpenRouter’s cache_control: ephemeral extension to the last user message (and last system message) so Anthropic / Gemini routes activate prompt caching. thinking is forwarded as reasoning_effort (low/medium/high) when set.

Parameters:
  • messages (list[Message]) – Conversation history.

  • model (str) – Model ID to use.

  • tools (list[ToolDefinition] | None) – Available tools for function calling.

  • temperature (float) – Sampling temperature.

  • max_tokens (int) – Maximum tokens in response.

  • tool_choice (ToolChoice | None) – How the model should select tools.

  • thinking (ThinkingConfig | None) – Extended thinking configuration. Forwarded as reasoning_effort when enabled.

  • enable_cache (bool) – Whether to enable prompt caching. When True, cache_control: ephemeral is attached to the last system and user message so OpenRouter’s Anthropic / Gemini backends activate caching.

Returns:

Tuple of (assistant message, tool calls if any).

Return type:

tuple[Message, list[ToolCall] | None]

Raises:

ProviderError – If not connected or request fails.

async chat_stream(messages, model, tools=None, temperature=0.7, max_tokens=4096, tool_choice=None, thinking=None, *, enable_cache=False)[source]

Stream a chat completion response from OpenRouter.

enable_cache attaches OpenRouter’s cache_control: ephemeral extension to the last user / system message so Anthropic and Gemini routes activate caching. thinking is forwarded as reasoning: {effort: ...} for backends that support it.

Parameters:
  • messages (list[Message]) – Conversation history.

  • model (str) – Model ID to use.

  • tools (list[ToolDefinition] | None) – Available tools for function calling.

  • temperature (float) – Sampling temperature.

  • max_tokens (int) – Maximum tokens in response.

  • tool_choice (ToolChoice | None) – How the model should select tools.

  • thinking (ThinkingConfig | None) – Extended thinking configuration. Forwarded as reasoning.effort when enabled.

  • enable_cache (bool) – Whether to enable prompt caching. Adds the cache_control: ephemeral extension to the last user / system message.

Yields:

str – Text chunks as they arrive.

Raises:
Return type:

AsyncIterator[str]

async cancel_request()[source]

Cancel any in-flight request.

Sets the cancel flag (which streaming loops poll) and cancels the active non-streaming task when one is registered, so both chat and chat_stream paths abort cleanly.

Return type:

None

async get_generation(generation_id)[source]

Get details about a specific generation.

Parameters:

generation_id (str) – The generation ID from a previous response.

Returns:

Generation details including cost and tokens used.

Return type:

dict[str, object]

Raises:

ProviderError – If not connected or request fails.