intellicrack.core.hexpat_compiler

HexPat DSL compiler that emits JSON templates from the shared AST.

This module is a thin compatibility/code-generation layer that delegates lexing and parsing to the canonical HexPat pipeline in intellicrack.core.hexpat and walks the resulting AST to produce a JSON template definition consumable by the Rust hex editor core.

The lexer, parser, AST, and error types are re-exported from this module for backward compatibility with previously published symbol paths.

class HexPatCodegen[source]

Bases: object

Generates JSON template definitions from a HexPat AST.

Walks the shared HexPat AST produced by intellicrack.core.hexpat.parser.HexPatParser and emits a JSON-serializable template definition. Runtime constructs that have no static representation are rejected during the walk.

__init__(declarations, pragma=None)[source]

Initialize the HexPatCodegen with parsed top-level AST nodes.

Top-level runtime-only declarations and statements raise HexPatError from _reject_runtime_top_level().

Parameters:
  • declarations (Sequence[DeclNode | StmtNode]) – Sequence of top-level AST nodes produced by HexPatParser. Items typically include StructDecl, UnionDecl, EnumDecl, and BitfieldDecl. Top-level placement statements or other StmtNode values are rejected at construction time as they have no representation in the static JSON template.

  • pragma (PragmaInfo | None) – Optional preprocessor-extracted #pragma metadata to propagate into the emitted JSON template. When supplied, the template’s default_endianness, description, author, and magic_detection fields are populated from the pragma values; otherwise the codegen falls back to inert defaults.

Return type:

None

generate()[source]

Generate the JSON template dict from all declarations.

Honours preprocessor-extracted #pragma metadata (when supplied at construction) by populating the template’s default_endianness, description, author, and magic_detection fields from the corresponding pragma values. #pragma base_address and #pragma bitfield_order are recorded under a pragma_metadata key so downstream consumers can apply them when materialising the template against binary data.

Returns:

JSON-serializable template definition.

Return type:

dict[str, Any]

Raises:

HexPatError – If no struct declaration is found.

class HexPatCompiler[source]

Bases: object

Compiles HexPat DSL source code into JSON template definitions.

Orchestrates the lexer, parser, and codegen pipeline. Lexing and parsing are delegated to the canonical implementations in intellicrack.core.hexpat; this class adds AST-walk validation and the static JSON code generator.

static compile(source)[source]

Compile DSL source to a JSON string.

On lexing, parsing, or code-generation failure, compile_to_dict() raises HexPatError.

Parameters:

source (str) – HexPat DSL source code.

Returns:

JSON template definition string.

Return type:

str

static compile_to_dict(source)[source]

Compile DSL source to a Python dict.

Runs the full preprocessor on source so that #pragma directives (endian, base_address, bitfield_order, author, description, magic, mime, pointer_size) are honoured by the generated static template. Without this step the codegen would silently fall back to inert defaults (little endian, generic description, no magic detection), which the audit identified as a codegen-behaviour drift.

Parameters:

source (str) – HexPat DSL source code.

Returns:

JSON-compatible template definition dict.

Return type:

dict[str, Any]

Raises:

HexPatError – If preprocessing, lexing, parsing, or code generation fails.

exception HexPatError[source]

Bases: Exception

Base error for the HexPat interpreter.

__init__(message, line=0, column=0, file='')[source]

Initialize the HexPatError with location and message details.

Parameters:
  • message (str) – Human-readable error description.

  • line (int) – Source line number where the error occurred.

  • column (int) – Source column number where the error occurred.

  • file (str) – Source file path where the error occurred.

Return type:

None

class HexPatLexer[source]

Bases: object

Tokenizer for the HexPat pattern language.

Converts raw source text into a flat list of Token objects. The final token in the returned list is always an EOF token.

__init__(source, file_path='<input>')[source]

Initialize the HexPatLexer with source text and file path.

Parameters:
  • source (str) – The raw source text to tokenize.

  • file_path (str) – Optional file path used in error messages.

Return type:

None

tokenize()[source]

Tokenize the source into a list of Tokens.

Returns:

A list of Token objects ending with an EOF token.

Return type:

list[Token]

class HexPatParser[source]

Bases: object

Recursive-descent parser for the HexPat .hexpat pattern language.

Consumes a flat list of tokens produced by the lexer and builds an AST consisting of top-level declarations and statements.

__init__(tokens, file_path='<input>')[source]

Initialize the HexPatParser with a token stream.

Parameters:
  • tokens (list[Token]) – Flat token list produced by the lexer.

  • file_path (str) – Source file path used for error location reporting.

Return type:

None

property errors: list[HexPatParseError]

All parse errors collected during recovery-mode parsing.

Returns:

Errors collected by parse(); empty if none.

Return type:

list[HexPatParseError]

parse()[source]

Parse the token stream into a list of top-level declarations and statements.

The parser attempts to recover from individual top-level syntax errors by collecting the raised HexPatParseError into errors and synchronising to the next top-level boundary (; or }). After the entire token stream has been processed, all collected errors are surfaced to the caller as a single HexPatAggregateParseError whose message enumerates every collected error and whose errors attribute holds the full list, so callers never silently lose information about secondary failures.

Returns:

Ordered list of top-level AST nodes parsed

from the token stream.

Return type:

list[DeclNode | StmtNode]

Raises:

HexPatAggregateParseError – When one or more syntax errors were collected. Subclass of HexPatParseError so existing except HexPatParseError clauses still catch it.

class Token[source]

Bases: object

A single lexer token.

Variables:
  • type (TokenType) – The token type.

  • value (str) – The raw source text of the token.

  • line (int) – Source line number (1-based).

  • column (int) – Source column number (1-based).

type: TokenType
value: str
line: int
column: int
__init__(type, value, line, column)
Parameters:
Return type:

None

class TokenType[source]

Bases: Enum

Token types produced by the HexPat lexer.

STRUCT = 'struct'
UNION = 'union'
ENUM = 'enum'
BITFIELD = 'bitfield'
IF = 'if'
ELSE = 'else'
MATCH = 'match'
WHILE = 'while'
FOR = 'for'
FN = 'fn'
RETURN = 'return'
BREAK = 'break'
CONTINUE = 'continue'
NAMESPACE = 'namespace'
USING = 'using'
IN = 'in'
TRY = 'try'
CATCH = 'catch'
U8 = 'u8'
U16 = 'u16'
U32 = 'u32'
U64 = 'u64'
U128 = 'u128'
S8 = 's8'
S16 = 's16'
S32 = 's32'
S64 = 's64'
S128 = 's128'
FLOAT = 'float'
DOUBLE = 'double'
CHAR = 'char'
CHAR16 = 'char16'
BOOL = 'bool'
STR = 'str'
AUTO = 'auto'
PADDING = 'padding'
LE = 'le'
BE = 'be'
SIZEOF = 'sizeof'
ADDRESSOF = 'addressof'
TYPENAMEOF = 'typenameof'
THIS = 'this'
PARENT = 'parent'
REF = 'ref'
OUT = 'out'
CONST = 'const'
NUMBER = 'number'
FLOAT_LITERAL = 'float_literal'
STRING_LITERAL = 'string_literal'
CHAR_LITERAL = 'char_literal'
TRUE_KW = 'true'
FALSE_KW = 'false'
NULL_KW = 'null'
IDENTIFIER = 'identifier'
LBRACE = '{'
RBRACE = '}'
LBRACKET = '['
RBRACKET = ']'
LPAREN = '('
RPAREN = ')'
SEMICOLON = ';'
COMMA = ','
DOT = '.'
COLON = ':'
DOUBLE_LBRACKET = '[['
DOUBLE_RBRACKET = ']]'
DOLLAR = '$'
AT = '@'
ARROW = '->'
DOUBLE_COLON = '::'
ELLIPSIS = '...'
EQ = '=='
NE = '!='
LT = '<'
GT = '>'
LE_OP = '<='
GE_OP = '>='
DOUBLE_AMPERSAND = '&&'
DOUBLE_PIPE = '||'
DOUBLE_CARET = '^^'
BANG = '!'
AMPERSAND = '&'
PIPE = '|'
CARET = '^'
TILDE = '~'
LSHIFT = '<<'
RSHIFT = '>>'
PLUS = '+'
MINUS = '-'
STAR = '*'
SLASH = '/'
PERCENT = '%'
ASSIGN = '='
PLUS_ASSIGN = '+='
MINUS_ASSIGN = '-='
STAR_ASSIGN = '*='
SLASH_ASSIGN = '/='
PERCENT_ASSIGN = '%='
AMPERSAND_ASSIGN = '&='
PIPE_ASSIGN = '|='
CARET_ASSIGN = '^='
LSHIFT_ASSIGN = '<<='
RSHIFT_ASSIGN = '>>='
QUESTION = '?'
EOF = 'eof'