intellicrack.core.script_gen

Script infrastructure for Intellicrack.

This module provides the stable Python API used by the application GUI layer (see main.py wiring) and by the tool/AI bridges to orchestrate AI-generated scripts. The actual script content is written dynamically by the AI based on analysis results - there are NO pre-built templates or generated scripts here.

Integration pattern:

A single ScriptGenerator instance is owned by the top-level application shell and passed (or re-referenced by composition) into GUI panels and bridge orchestrators that need to build AI prompts from ScriptContext state. ScriptManager owns the on-disk storage and in-memory cache of Script objects, and ScriptValidator is used to validate script contents before handing them to an external execution bridge.

The AI creates scripts from scratch using: - Analysis results from analysis_aggregator - Binary metadata from the disassembly tools - Runtime information from Frida/debugger sessions

This module only provides: - Data classes for script metadata and storage - Validation utilities for script syntax - Execution context information - Script management (save, load, execute)

strip_java_strings_and_comments(content)[source]

Strip Java string literals, char literals, and comments from source.

Replaces every string literal, character literal, line comment, and block comment with whitespace of the same length. Preserves newline characters so line numbers remain stable for downstream analysis.

Parameters:

content (str) – Java source text to scrub.

Returns:

Sanitised source where only true Java tokens are visible to keyword/brace counting routines.

Return type:

str

class ScriptLanguage[source]

Bases: Enum

Script language enumeration.

JAVASCRIPT = 'javascript'
JAVA = 'java'
PYTHON = 'python'
R2_COMMANDS = 'r2_commands'
X64DBG_SCRIPT = 'x64dbg_script'
class BypassStrategy[source]

Bases: Enum

Bypass strategy types for license protections.

These are hints for the AI when writing scripts, not template selectors.

Variables:
  • RETURN_TRUE – Force target function to return true (1).

  • RETURN_FALSE – Force target function to return false (0).

  • RETURN_ZERO – Force target function to return integer zero.

  • RETURN_ONE – Force target function to return integer one.

  • NOP_FUNCTION – Replace target function body with a no-op.

  • SKIP_CHECK – Skip over a conditional check entirely.

  • PATCH_JUMP – Patch a conditional jump to always or never branch.

  • HOOK_REPLACE – Replace entire function with a custom implementation.

  • MEMORY_PATCH – Patch bytes in process memory at runtime.

  • INLINE_PATCH – Patch instruction bytes directly in the binary on disk.

  • VIRTUALIZATION_DEFEAT – Bypass code-virtualization protections (VMProtect, Themida, etc.).

RETURN_TRUE = 'return_true'
RETURN_FALSE = 'return_false'
RETURN_ZERO = 'return_zero'
RETURN_ONE = 'return_one'
NOP_FUNCTION = 'nop_function'
SKIP_CHECK = 'skip_check'
PATCH_JUMP = 'patch_jump'
HOOK_REPLACE = 'hook_replace'
MEMORY_PATCH = 'memory_patch'
INLINE_PATCH = 'inline_patch'
VIRTUALIZATION_DEFEAT = 'virtualization_defeat'
property description: str

Human-readable description of the strategy.

Returns:

Description text for this bypass strategy.

Return type:

str

class ScriptContext[source]

Bases: object

Context information for AI script generation.

Provides the AI with all necessary information to write an effective script.

Variables:
  • binary_name (str) – Name of the binary being analyzed.

  • binary_path (Path | None) – Full path to the binary.

  • architecture (str) – Target architecture (x86, x64, arm, arm64).

  • platform (str) – Target platform (windows, linux, macos).

  • module_base (int | None) – Base address of the main module (if known).

  • target_functions (list[dict[str, Any]]) – Functions identified for bypass/hooking.

  • identified_protections (list[str]) – Protection mechanisms detected.

  • crypto_apis (list[str]) – Crypto API calls found in the binary.

  • string_references (list[str]) – Relevant string references found.

  • magic_constants (list[int]) – Magic constants used in validation.

  • additional_context (dict[str, Any]) – Any additional context from analysis.

  • LANGUAGE_API_MAP (ClassVar[dict[ScriptLanguage, str]]) – Class-level mapping from ScriptLanguage to the API-reference key ("frida", "ghidra", "cutter", "x64dbg") used to look up the per-language reference dict.

binary_name: str
binary_path: Path | None = None
architecture: str = 'x64'
platform: str = 'windows'
module_base: int | None = None
target_functions: list[dict[str, Any]]
identified_protections: list[str]
crypto_apis: list[str]
string_references: list[str]
magic_constants: list[int]
additional_context: dict[str, Any]
LANGUAGE_API_MAP: ClassVar[dict[ScriptLanguage, str]] = {ScriptLanguage.JAVA: 'ghidra', ScriptLanguage.JAVASCRIPT: 'frida', ScriptLanguage.R2_COMMANDS: 'cutter', ScriptLanguage.X64DBG_SCRIPT: 'x64dbg'}
to_prompt_context(language=None)[source]

Convert context to a string suitable for AI prompts.

Parameters:

language (ScriptLanguage | None) – Target script language (optional) to include API reference.

Returns:

Formatted context string.

Return type:

str

__init__(binary_name, binary_path=None, architecture='x64', platform='windows', module_base=None, target_functions=<factory>, identified_protections=<factory>, crypto_apis=<factory>, string_references=<factory>, magic_constants=<factory>, additional_context=<factory>)
Parameters:
Return type:

None

class Script[source]

Bases: object

A script ready for execution.

Variables:
  • name (str) – Script name identifier.

  • script_type (Literal['frida', 'ghidra', 'cutter', 'python', 'x64dbg']) – Type of the script.

  • language (ScriptLanguage) – Programming language of the script.

  • content (str) – Script source code content.

  • description (str) – Human-readable description of the script.

  • created_at (datetime) – Generation timestamp (tz-aware UTC).

  • context (ScriptContext | None) – Context used to generate the script.

  • target_functions (list[str]) – Target functions the script operates on.

  • verified (bool) – Whether the script has been syntax-verified.

  • execution_results (dict[str, Any]) – Results from script execution (if run).

  • saved_path (Path | None) – On-disk path of the most recent successful save.

name: str
script_type: Literal['frida', 'ghidra', 'cutter', 'python', 'x64dbg']
language: ScriptLanguage
content: str
description: str
created_at: datetime
context: ScriptContext | None = None
target_functions: list[str]
verified: bool = False
execution_results: dict[str, Any]
saved_path: Path | None = None
add_execution_result(tool_name, result)[source]

Add or update an execution result record.

Parameters:
  • tool_name (str) – Name of the tool that executed the script.

  • result (object) – The result object or data.

Return type:

None

save(path)[source]

Save script to file.

Writes the script’s content to path after ensuring the parent directory exists. The success log record is emitted only after the write returns successfully; if the write raises an OSError, the failure is logged with the exception detail and re-raised so callers can react.

Parameters:

path (Path) – File path to save to.

Raises:

OSError – If the parent directory cannot be created or the write_text call fails for any I/O reason.

Return type:

None

get_extension()[source]

Get the appropriate file extension for this script type.

Returns:

File extension including the dot.

Return type:

str

__init__(name, script_type, language, content, description, created_at=<factory>, context=None, target_functions=<factory>, verified=False, execution_results=<factory>, saved_path=None)
Parameters:
Return type:

None

class ScriptValidator[source]

Bases: object

Validates script syntax before execution.

static validate_python(content)[source]

Validate Python script syntax.

Parameters:

content (str) – Python script content.

Returns:

Tuple of (is_valid, error_message).

Return type:

tuple[bool, str | None]

static validate_javascript(content)[source]

Validate JavaScript syntax using the node runtime.

Writes the script to a temporary file and invokes node --check to perform real syntax validation. A (True, None) return is only produced when node actually reports a clean parse (exit code 0).

Parameters:

content (str) – JavaScript script content.

Returns:

(True, None) when node confirms the script parses cleanly. (False, <reason>) in every other case, including node being unavailable, tempfile write failures, subprocess timeouts, and actual syntax errors. The second element is always populated when the first is False.

Return type:

tuple[bool, str | None]

static validate_java(content)[source]

Validate Java/Ghidra script structure.

Performs a structural check on a Ghidra Java script. String literals, character literals, line comments, and block comments are stripped from the source before any keyword or brace check runs, so tokens appearing inside strings or comments cannot satisfy the structural requirements and braces inside string literals do not contribute to the brace-balance count.

Parameters:

content (str) – Java script content.

Returns:

(True, None) when the sanitised script contains an import statement, a public modifier, an explicit class declaration, and a void run( entry point with balanced braces. (False, <reason>) otherwise.

Return type:

tuple[bool, str | None]

validate(script)[source]

Validate a script based on its language.

A script whose content is empty or only whitespace is rejected with (False, "Script is empty") before any language validator runs. This closes the case where an empty program parses cleanly (an empty string is valid Python and valid JavaScript), which would otherwise report success when no script is actually present.

Languages without a real validator (currently R2_COMMANDS and X64DBG_SCRIPT) return (False, <reason>) and leave Script.verified unchanged. The previous behaviour of silently returning success is removed because callers used the verified flag to gate execution and would otherwise treat unvalidated scripts as trusted.

Parameters:

script (Script) – Script to validate.

Returns:

(True, None) when the language has a validator and the non-empty script passes. (False, <reason>) when the script is empty, fails validation, or no validator exists for the language.

Return type:

tuple[bool, str | None]

class ScriptManager[source]

Bases: object

Manages script storage, retrieval, and execution.

Variables:
  • scripts_dir (Path) – Directory for storing scripts.

  • scripts (dict[str, Script]) – Mapping of script names to Script objects.

__init__(scripts_dir=None)[source]

Initialize the ScriptManager with a storage directory.

Parameters:

scripts_dir (Path | None) – Directory for storing scripts. When omitted, a scripts directory beneath the current working directory is used. The directory is not created on the filesystem until a script is actually written.

Return type:

None

scripts_dir: Path
scripts: dict[str, Script]
add_script(script, *, validate=True)[source]

Add a script to the manager.

Parameters:
  • script (Script) – Script to add.

  • validate (bool) – Whether to validate syntax before adding.

Returns:

True if script was added successfully.

Return type:

bool

get_script(name)[source]

Get a script by name.

Parameters:

name (str) – Script name.

Returns:

Script or None if not found.

Return type:

Script | None

delete_script(name)[source]

Delete a script by name, removing both memory and disk state.

Removes the in-memory entry and, when the script has a backing file on disk, unlinks it so a deleted script cannot be reloaded from a stale file. The backing path is derived with _resolve_saved_path(), the same logic save_script() and reload_script() use.

If the file exists but cannot be removed (permission denied, the file is held open by another process, or any other OS-level failure), the failure is logged and the in-memory entry is retained so the caller can retry the delete rather than being left in a state where the script vanished from the list but its file survived.

Parameters:

name (str) – Script name to delete.

Returns:

True if the script was deleted, meaning the in-memory entry was removed and any backing file it had was unlinked or never existed. False if no script named name is registered, or if a backing file exists but could not be removed.

Return type:

bool

list_scripts(script_type=None)[source]

List available scripts.

Parameters:

script_type (Literal['frida', 'ghidra', 'cutter', 'python', 'x64dbg'] | None) – Optional filter by script type.

Returns:

List of script names.

Return type:

list[str]

save_script(name, subdir=None)[source]

Save a script to disk.

Parameters:
  • name (str) – Script name.

  • subdir (str | None) – Optional subdirectory within scripts_dir.

Returns:

Path where script was saved, or None if not found.

Return type:

Path | None

load_script(path)[source]

Load a script from disk.

Parameters:

path (Path) – Path to script file.

Returns:

Loaded script or None if failed.

Return type:

Script | None

ensure_script_saved(name)[source]

Ensure a script is saved to disk.

Parameters:

name (str) – Script name.

Returns:

True if saved successfully.

Return type:

bool

reload_script(name)[source]

Reload a script from disk.

When save_script() was previously called for name, the recorded Script.saved_path is used directly so reloads survive subdirectory writes. When no save path is recorded the manager falls back to the canonical scripts_dir / <name><ext> location.

Parameters:

name (str) – Script name.

Returns:

True if reloaded successfully.

Return type:

bool

record_execution(script_name, tool_name, result)[source]

Record an execution result for a script.

Forwards to the script’s add_execution_result method to persist tool execution metadata.

Parameters:
  • script_name (str) – Name of the script that was executed.

  • tool_name (str) – Name of the tool that executed the script.

  • result (object) – The result object or data from execution.

Returns:

True if the result was recorded, False if script not found.

Return type:

bool

execute(name, *, args=None, timeout=30.0, cwd=None)[source]

Execute a script via the language-appropriate runner.

Dispatches the script’s content to the right interpreter based on Script.language. Every supported language has a real runner wired up so the integration is end-to-end:

  • PYTHON is run with the active interpreter (sys.executable).

  • JAVASCRIPT is run with node.

  • R2_COMMANDS are piped to r2 via -q -i.

  • JAVA (Ghidra script) is launched with Ghidra’s analyzeHeadless driver and -postScript.

  • X64DBG_SCRIPT is run with the platform-appropriate x64dbg binary using its -script switch.

Scripts that have not yet been saved are written to a fresh temporary file before invocation so the runner sees a real path. The execution is tracked through ProcessManager, the exit code and captured stdout/stderr are recorded on the script via record_execution(), and the underlying subprocess.CompletedProcess is returned to the caller so no diagnostic information is lost.

Parameters:
  • name (str) – Name of the script to execute.

  • args (list[str] | None) – Optional argument list to forward to the runner after the script path.

  • timeout (float | None) – Maximum seconds to wait for the runner to exit. Pass None to disable the timeout.

  • cwd (Path | None) – Optional working directory to use for the runner.

Returns:

The completed subprocess with text stdout/stderr and exit code.

Return type:

CompletedProcess[str]

Raises:
  • KeyError – If no script with name is registered.

  • FileNotFoundError – If the runner binary required by the script’s language cannot be located on PATH.

  • TimeoutExpired – If the runner does not exit within timeout seconds.

build_execute_command(script, args)[source]

Build the runner command line for a script.

Materialises script to disk via Script.saved_path if already persisted, otherwise writes the content to a temporary file with the right extension. Returns the runner argv for the script’s language.

Parameters:
  • script (Script) – Script to execute.

  • args (list[str] | None) – Optional argument list to append after the script path.

Returns:

Runner command line ready to pass to ProcessManager.run_tracked().

Return type:

list[str]

get_frida_api_reference()[source]

Get Frida API reference for AI context.

Returns:

Dictionary mapping API categories to usage examples.

Return type:

dict[str, str]

get_ghidra_api_reference()[source]

Get Ghidra API reference for AI context.

Returns:

Dictionary mapping API categories to usage examples.

Return type:

dict[str, str]

get_cutter_reference()[source]

Get Cutter/Rizin command reference for AI context.

Returns:

Dictionary mapping command categories to examples.

Return type:

dict[str, str]

get_x64dbg_reference()[source]

Get x64dbg command reference for AI context.

Returns:

Dictionary mapping command categories to examples.

Return type:

dict[str, str]

class ScriptGenerator[source]

Bases: object

Stateful entry point for building AI prompts that generate scripts.

ScriptGenerator is the API surface consumed by the Intellicrack application shell (main.py) and by tool/AI bridges that need to turn a ScriptContext into a prompt string ready for a language model.

The instance owns three pieces of state that are referenced by the public generate_* helpers:

  • validator - a ScriptValidator instance used to pre-flight script content before it is shipped to an external execution bridge.

  • output_dir - the directory under which generated scripts and drafts are persisted. Defaults to Path.cwd() / "generated_scripts" and is created lazily by prepare_output_path().

  • An API-reference cache populated lazily by api_reference() so the language-specific reference dicts are computed at most once per (language, generator) pair.

All constructor parameters are optional so existing call sites (main.py, ui/app.py, ui/tools.py, ui/panels/script_manager.py, core/orchestrator.py) keep compiling with ScriptGenerator().

DEFAULT_OUTPUT_DIRNAME: ClassVar[str] = 'generated_scripts'
__init__(validator=None, output_dir=None)[source]

Initialize the ScriptGenerator with stateful dependencies.

Parameters:
  • validator (ScriptValidator | None) – Optional ScriptValidator instance to reuse for pre-flight validation. A fresh instance is created when not supplied.

  • output_dir (Path | None) – Optional directory under which generated scripts and drafts will be persisted. Defaults to Path.cwd() / "generated_scripts". The directory is materialised on the filesystem only when prepare_output_path() is invoked.

Return type:

None

validator: ScriptValidator
output_dir: Path
api_reference(language)[source]

Return the cached API reference for language.

Looks up the language-specific reference dict, populating the per-instance cache on first request. Languages without a known reference (currently PYTHON) return an empty dict.

Parameters:

language (ScriptLanguage) – Script language whose reference to fetch.

Returns:

Reference categories mapped to usage examples, or an empty dict when no reference is registered for the language.

Return type:

dict[str, str]

prepare_output_path(name, language)[source]

Resolve and ensure the output path for a generated script.

Materialises output_dir on the filesystem (creating any missing parents) and returns the canonical path that name/language should be written to. The file itself is not created here; only the directory is ensured.

Parameters:
  • name (str) – Script base name without extension.

  • language (ScriptLanguage) – Target script language; controls the file extension.

Returns:

Absolute path the caller can pass to Script.save() or ScriptManager.save_script().

Return type:

Path

prepare_ai_prompt(language)

Build the AI prompt, dispatching as bound or unbound automatically.

Calling generator.prepare_ai_prompt(context, language) consults the instance’s API-reference cache. Calling ScriptGenerator.prepare_ai_prompt(context, language) keeps legacy free-function-style call sites working by resolving the API reference fresh on each invocation. The signature in both cases is (context: ScriptContext, language: ScriptLanguage) -> str.

Parameters:
Return type:

str

generate_frida(context)[source]

Build an AI prompt targeted at Frida (JavaScript) script generation.

Parameters:

context (ScriptContext) – Analysis context to embed in the prompt.

Returns:

Prompt text with the Frida/JavaScript API reference included.

Return type:

str

generate_ghidra(context)[source]

Build an AI prompt targeted at Ghidra (Java) script generation.

Parameters:

context (ScriptContext) – Analysis context to embed in the prompt.

Returns:

Prompt text with the Ghidra Java API reference included.

Return type:

str

generate_python(context)[source]

Build an AI prompt targeted at generic Python script generation.

Parameters:

context (ScriptContext) – Analysis context to embed in the prompt.

Returns:

Prompt text for a Python script (no vendor API reference).

Return type:

str

generate_cutter(context)[source]

Build an AI prompt targeted at Cutter/Rizin command script generation.

Parameters:

context (ScriptContext) – Analysis context to embed in the prompt.

Returns:

Prompt text with the Cutter/Rizin command reference included.

Return type:

str

generate_x64dbg(context)[source]

Build an AI prompt targeted at x64dbg script generation.

Parameters:

context (ScriptContext) – Analysis context to embed in the prompt.

Returns:

Prompt text with the x64dbg command reference included.

Return type:

str