Building a Context-Aware Parsing Engine for Mobile Architectures
Dumping an entire multi-module mobile repository into a large context window does not work. While modern models advertise context windows of one million tokens or more, retrieval accuracy degrades rapidly when the input is unstructured raw text. In mobile development, where architectures rely heavily on implicit dependency injection, protocol-oriented programming, and multi-module separation, raw text search fails to capture the execution path.
To solve this, we built a context-aware parsing engine. Instead of treating a codebase as a flat sequence of characters, the engine maps the repository into a deterministic directed graph. This graph represents the structural, logical, and architectural relationships of the codebase. When the planner needs to reason about a specific component, the engine queries this graph to extract a highly relevant, pruned subgraph. This approach ensures that the context window contains only the precise code paths and interfaces necessary for the task, eliminating noise and preserving model attention.
The Structural Complexity of Mobile Codebases
Mobile applications present unique challenges for static analysis. Unlike monolithic backend systems that often follow linear execution paths, modern iOS and Android applications are highly decoupled. They rely on:
- Protocol-oriented programming or interface-driven development.
- Dependency injection frameworks that resolve implementations at runtime.
- Multi-module architectures where code is split across dozens of independent targets.
If a developer asks the system to modify a payment flow, a naive search might locate the view controller. However, the actual business logic resides in a presenter or view model located in a separate module, which communicates via a protocol defined in a third API module. The concrete implementation of that protocol is injected at runtime by a dependency injection container.
Without a deterministic map of these relationships, the system cannot locate the code that actually needs to change. It is forced to guess, leading to incomplete context and broken implementations.
Designing the Context-Aware Parsing Engine
The engine operates in three distinct phases: extraction, resolution, and pruning.
Phase 1: Extraction
We use tree-sitter parsers to generate concrete syntax trees for every source file in the repository. From these trees, we extract key architectural entities:
- Classes, structs, and actors
- Protocols and interfaces
- Method signatures and properties
- Import statements and module declarations
- Dependency injection registrations
Phase 2: Resolution
Once the entities are extracted, the engine resolves the relationships between them. This is where we build the directed graph. Nodes represent entities, and edges represent relationships such as inheritance, conformance, imports, or calls.
Resolving implicit dependencies requires custom heuristics. For example, if class A's initializer accepts an argument of protocol type P, the engine searches the graph for classes that conform to P. If a dependency injection module binds P to concrete class B, the engine creates a directed edge from A to B labeled as a dependency.
Phase 3: Pruning
When a task is initiated, the engine does not send the entire graph. Instead, it identifies the target file and performs a depth-limited traversal of the graph. It selects the target file, its immediate dependencies, the protocols it conforms to, and the concrete implementations of those protocols. This pruned subgraph is then serialized into a structured format for the planner.
Implementation Sketch
Below is a pseudo-code implementation demonstrating how the parsing engine processes a file, extracts entities, and registers them into a dependency graph.
# Pseudo-code: Context-Aware Graph Construction Engine
class Entity:
def __init__(self, name: str, entity_type: str, file_path: str):
self.name = name
self.entity_type = entity_type # e.g., "class", "protocol", "struct"
self.file_path = file_path
self.dependencies = []
self.conformances = []
class CodebaseGraph:
def __init__(self):
self.nodes = {} # Map of entity name to Entity object
self.edges = [] # List of tuples (source, target, relation_type)
def add_entity(self, entity: Entity):
self.nodes[entity.name] = entity
def add_relation(self, source: str, target: str, relation_type: str):
self.edges.append((source, target, relation_type))
def resolve_protocols(self):
# Resolve which concrete classes implement which protocols
for node_name, entity in self.nodes.items():
for conformance in entity.conformances:
if conformance in self.nodes:
self.add_relation(node_name, conformance, "conforms_to")
class ParsingEngine:
def __init__(self, graph: CodebaseGraph):
self.graph = graph
def parse_file(self, file_path: str, file_content: str):
# In practice, this uses tree-sitter to walk the AST
# Here we simulate the extraction of a class and its dependencies
lines = file_content.split("\n")
current_entity = None
for line in lines:
line = line.strip()
if line.startswith("class ") or line.startswith("struct "):
parts = line.split()
name = parts[1].split(":")[0].strip()
current_entity = Entity(name, "class", file_path)
# Check for inheritance or protocol conformance
if ":" in line:
conformance_part = line.split(":")[1]
conformances = [c.strip() for c in conformance_part.split(",")]
current_entity.conformances.extend(conformances)
self.graph.add_entity(current_entity)
elif "init(" in line or "constructor(" in line:
# Extract injected dependencies from initializer signature
if current_entity:
# Simple heuristic to find typed parameters
# e.g., init(service: PaymentService)
if ":" in line:
param_part = line.split("(")[1].split(")")[0]
for param in param_part.split(","):
if ":" in param:
dep_type = param.split(":")[1].strip()
current_entity.dependencies.append(dep_type)
def finalize_graph(self):
# Connect dependencies to their resolved entities
for node_name, entity in self.graph.nodes.items():
for dep in entity.dependencies:
if dep in self.graph.nodes:
self.graph.add_relation(node_name, dep, "depends_on")
self.graph.resolve_protocols()
Managing Multi-Module Boundaries
In large mobile applications, code is frequently split across dozens of local modules or packages to improve build times and enforce separation of concerns. A change in a core module can have cascading effects on multiple feature modules.
Our parsing engine handles this by parsing package manifests, such as Swift Package Manager Package.swift files or Gradle build files. By mapping module-level dependencies first, the engine establishes a high-level topology of the repository. When a class in a feature module is modified, the engine uses this topology to identify which upstream modules must be analyzed for potential breaking changes, and which downstream modules must be updated to accommodate the new interface.
Evaluating the Trade-offs
Building a custom parsing engine involves significant trade-offs compared to using off-the-shelf vector databases or simple keyword search.
1. Maintenance Overhead vs. Precision
Vector databases are language-agnostic and require zero configuration. However, they lack semantic understanding of code relationships. A vector search for a specific service might return the UI layout file, the network client, and the local database schema, but it cannot tell the planner how these files interact. The parsing engine requires custom AST queries for each language, but it provides absolute precision.
2. Static vs. Dynamic Analysis
Our engine relies entirely on static analysis. While dynamic analysis, such as tracing execution paths at runtime, can capture highly dynamic behaviors, it requires compiling and running the application. In mobile development, compilation is slow and environment-dependent. Static analysis allows the engine to map the codebase instantly without requiring a successful build environment.
3. Graph Density vs. Context Window Limits
If the graph is too dense, traversing it can pull in too many files, defeating the purpose of pruning. We address this by categorizing relationships. Imports are treated as weak edges, while inheritance and dependency injection bindings are treated as strong edges. The pruning algorithm prioritizes strong edges, ensuring that only the critical execution path is loaded into the context window.
The Fallacy of the Infinite Context Window
There is a common belief that as context windows expand, the need for sophisticated retrieval systems diminishes. This is a misunderstanding of how transformer models process information. Attention mechanisms scale quadratically with sequence length. More importantly, irrelevant code acts as distractor material, increasing the likelihood of hallucinations or missed details.
By front-loading the structural analysis into a deterministic parsing engine, we narrow the decision space before the planner ever receives the prompt. The planner does not need to figure out how the modules are connected; the engine has already solved that. This separation of concerns, deterministic mapping for structure, probabilistic reasoning for implementation, is the only reliable way to handle large-scale mobile architectures.
Conclusion
Mapping complex mobile codebases into deterministic graphs is essential for accurate code generation and analysis. By combining AST parsing with dependency resolution, a context-aware parsing engine ensures that the planner operates with the exact structural context it needs.
To learn more about how we design deterministic systems for mobile development, read the full article on bridgedev.io: https://bridgedev.io/blog/building-our-context-aware-parsing-engine-for-mobile?utm_source=devto&utm_medium=social&utm_campaign=blog
Top comments (0)