intellicrack.core.hexpat_compiler
HexPat DSL compiler that emits JSON templates from the shared AST.
This module is a thin compatibility/code-generation layer that delegates
lexing and parsing to the canonical HexPat pipeline in
intellicrack.core.hexpat and walks the resulting AST to produce a
JSON template definition consumable by the Rust hex editor core.
The lexer, parser, AST, and error types are re-exported from this module for backward compatibility with previously published symbol paths.
- class HexPatCodegen[source]
Bases:
objectGenerates JSON template definitions from a HexPat AST.
Walks the shared HexPat AST produced by
intellicrack.core.hexpat.parser.HexPatParserand emits a JSON-serializable template definition. Runtime constructs that have no static representation are rejected during the walk.- __init__(declarations, pragma=None)[source]
Initialize the HexPatCodegen with parsed top-level AST nodes.
Top-level runtime-only declarations and statements raise
HexPatErrorfrom_reject_runtime_top_level().- Parameters:
declarations (Sequence[DeclNode | StmtNode]) – Sequence of top-level AST nodes produced by
HexPatParser. Items typically includeStructDecl,UnionDecl,EnumDecl, andBitfieldDecl. Top-level placement statements or otherStmtNodevalues are rejected at construction time as they have no representation in the static JSON template.pragma (PragmaInfo | None) – Optional preprocessor-extracted
#pragmametadata to propagate into the emitted JSON template. When supplied, the template’sdefault_endianness,description,author, andmagic_detectionfields are populated from the pragma values; otherwise the codegen falls back to inert defaults.
- Return type:
None
- generate()[source]
Generate the JSON template dict from all declarations.
Honours preprocessor-extracted
#pragmametadata (when supplied at construction) by populating the template’sdefault_endianness,description,author, andmagic_detectionfields from the corresponding pragma values.#pragma base_addressand#pragma bitfield_orderare recorded under apragma_metadatakey so downstream consumers can apply them when materialising the template against binary data.- Returns:
JSON-serializable template definition.
- Return type:
- Raises:
HexPatError – If no struct declaration is found.
- class HexPatCompiler[source]
Bases:
objectCompiles HexPat DSL source code into JSON template definitions.
Orchestrates the lexer, parser, and codegen pipeline. Lexing and parsing are delegated to the canonical implementations in
intellicrack.core.hexpat; this class adds AST-walk validation and the static JSON code generator.- static compile(source)[source]
Compile DSL source to a JSON string.
On lexing, parsing, or code-generation failure,
compile_to_dict()raisesHexPatError.
- static compile_to_dict(source)[source]
Compile DSL source to a Python dict.
Runs the full preprocessor on
sourceso that#pragmadirectives (endian,base_address,bitfield_order,author,description,magic,mime,pointer_size) are honoured by the generated static template. Without this step the codegen would silently fall back to inert defaults (littleendian, generic description, no magic detection), which the audit identified as a codegen-behaviour drift.- Parameters:
source (str) – HexPat DSL source code.
- Returns:
JSON-compatible template definition dict.
- Return type:
- Raises:
HexPatError – If preprocessing, lexing, parsing, or code generation fails.
- exception HexPatError[source]
Bases:
ExceptionBase error for the HexPat interpreter.
- class HexPatLexer[source]
Bases:
objectTokenizer for the HexPat pattern language.
Converts raw source text into a flat list of Token objects. The final token in the returned list is always an EOF token.
- class HexPatParser[source]
Bases:
objectRecursive-descent parser for the HexPat .hexpat pattern language.
Consumes a flat list of tokens produced by the lexer and builds an AST consisting of top-level declarations and statements.
- property errors: list[HexPatParseError]
All parse errors collected during recovery-mode parsing.
- Returns:
Errors collected by
parse(); empty if none.- Return type:
- parse()[source]
Parse the token stream into a list of top-level declarations and statements.
The parser attempts to recover from individual top-level syntax errors by collecting the raised
HexPatParseErrorintoerrorsand synchronising to the next top-level boundary (;or}). After the entire token stream has been processed, all collected errors are surfaced to the caller as a singleHexPatAggregateParseErrorwhose message enumerates every collected error and whoseerrorsattribute holds the full list, so callers never silently lose information about secondary failures.- Returns:
- Ordered list of top-level AST nodes parsed
from the token stream.
- Return type:
- Raises:
HexPatAggregateParseError – When one or more syntax errors were collected. Subclass of
HexPatParseErrorso existingexcept HexPatParseErrorclauses still catch it.
- class TokenType[source]
Bases:
EnumToken types produced by the HexPat lexer.
- STRUCT = 'struct'
- UNION = 'union'
- ENUM = 'enum'
- BITFIELD = 'bitfield'
- IF = 'if'
- ELSE = 'else'
- MATCH = 'match'
- WHILE = 'while'
- FOR = 'for'
- FN = 'fn'
- RETURN = 'return'
- BREAK = 'break'
- CONTINUE = 'continue'
- NAMESPACE = 'namespace'
- USING = 'using'
- IN = 'in'
- TRY = 'try'
- CATCH = 'catch'
- U8 = 'u8'
- U16 = 'u16'
- U32 = 'u32'
- U64 = 'u64'
- U128 = 'u128'
- S8 = 's8'
- S16 = 's16'
- S32 = 's32'
- S64 = 's64'
- S128 = 's128'
- FLOAT = 'float'
- DOUBLE = 'double'
- CHAR = 'char'
- CHAR16 = 'char16'
- BOOL = 'bool'
- STR = 'str'
- AUTO = 'auto'
- PADDING = 'padding'
- LE = 'le'
- BE = 'be'
- SIZEOF = 'sizeof'
- ADDRESSOF = 'addressof'
- TYPENAMEOF = 'typenameof'
- THIS = 'this'
- PARENT = 'parent'
- REF = 'ref'
- OUT = 'out'
- CONST = 'const'
- NUMBER = 'number'
- FLOAT_LITERAL = 'float_literal'
- STRING_LITERAL = 'string_literal'
- CHAR_LITERAL = 'char_literal'
- TRUE_KW = 'true'
- FALSE_KW = 'false'
- NULL_KW = 'null'
- IDENTIFIER = 'identifier'
- LBRACE = '{'
- RBRACE = '}'
- LBRACKET = '['
- RBRACKET = ']'
- LPAREN = '('
- RPAREN = ')'
- SEMICOLON = ';'
- COMMA = ','
- DOT = '.'
- COLON = ':'
- DOUBLE_LBRACKET = '[['
- DOUBLE_RBRACKET = ']]'
- DOLLAR = '$'
- AT = '@'
- ARROW = '->'
- DOUBLE_COLON = '::'
- ELLIPSIS = '...'
- EQ = '=='
- NE = '!='
- LT = '<'
- GT = '>'
- LE_OP = '<='
- GE_OP = '>='
- DOUBLE_AMPERSAND = '&&'
- DOUBLE_PIPE = '||'
- DOUBLE_CARET = '^^'
- BANG = '!'
- AMPERSAND = '&'
- PIPE = '|'
- CARET = '^'
- TILDE = '~'
- LSHIFT = '<<'
- RSHIFT = '>>'
- PLUS = '+'
- MINUS = '-'
- STAR = '*'
- SLASH = '/'
- PERCENT = '%'
- ASSIGN = '='
- PLUS_ASSIGN = '+='
- MINUS_ASSIGN = '-='
- STAR_ASSIGN = '*='
- SLASH_ASSIGN = '/='
- PERCENT_ASSIGN = '%='
- AMPERSAND_ASSIGN = '&='
- PIPE_ASSIGN = '|='
- CARET_ASSIGN = '^='
- LSHIFT_ASSIGN = '<<='
- RSHIFT_ASSIGN = '>>='
- QUESTION = '?'
- EOF = 'eof'