Skip to content

Architecture Analyzer

Derive a codebase's architectural principles as deterministic, checkable facts rather than inferred opinions.

Answers: "How is this built, and on what principles?" by analyzing the import graph with no inference, no model, and no guessing from names. Every verdict is arithmetic over the module dependency graph and the symbol table, and every verdict ships the evidence it was computed from — module lists, counts, and ratios — so a reader can verify rather than trust.

Usage

CLI

drydock architecture <project_path>
drydock architecture <project_path> --output arch.json
drydock architecture <project_path> --markdown

Default output is JSON. Add --markdown or -m for human-readable format.

MCP

drydock_architecture(project_path: str) -> str

Returns architecture analysis as JSON by default.

The Point: Deterministic, Not Inferred

This analyzer is arithmetic. It computes metrics over the actual import graph and symbol table — the same inputs as a compiler. Identical input always produces identical output, and a reader can verify every claim by inspecting the evidence.

This is what makes it useful for learning from unfamiliar codebases. An architectural opinion unsupported by checkable data teaches wrong patterns confidently. An architectural fact with its evidence teaches correctly.

Example: Dependency Inversion

Bad output (opinion): "Uses dependency inversion."

This analyzer (fact): "2 of the 10 most-depended-upon packages define an abstract contract (Protocol, ABC, trait or interface). Here they are: pkg.base, pkg.models."

The reader can verify: Are those really the hubs? Do they define contracts? The data is there.

The 8 Principles

Every principle answers one architectural question. All are deterministic; none consult a model.

ID Question What Evidence Returns
Acyclic Dependencies Can modules load in some order without circularity? Count of load-time cycles, deferred-only cycles, and edge counts
Layering Do dependencies flow in one direction through distinct levels? Number of levels, modules per level, package alignment
Stable Dependencies Do modules depend toward greater stability? Count and ratio of violating edges (instable→stable)
Stable Abstractions Are widely-depended-upon modules abstract rather than concrete? Module counts in zone of pain and zone of uselessness, mean distance
Dependency Inversion Do the most-depended-upon modules define contracts? List of hub packages with their abstract/concrete type counts
No God Modules Is any single module a bottleneck in both directions? Threshold, modules exceeding it, widest fan-in/fan-out
Tests Target Public Surface Do tests exercise public entry points, or reach into internals? Test module count, private import count, violations
Plugin Seams Are there contracts with implementations that could be lifted out whole? List of seams ranked by cohesion, implementer counts, escaping dependencies

Key Metrics

All metrics are computed per module (file), though some principles aggregate to packages.

Stability & Distance Metrics (Martin)

These are standard package-design metrics, adapted to the module level:

  • Ca (afferent coupling): how many modules depend on this one
  • Ce (efferent coupling): how many modules this one depends on
  • I (instability): Ce / (Ca + Ce) where 0 = maximally stable, 1 = maximally unstable
  • A (abstractness): (abstract types) / (all types) where abstract means the language's native contract marker (Protocol, ABC, interface, trait)
  • D (distance from main sequence): |A + I - 1| where 0 is healthy (either stable-abstract or unstable-concrete) and ~1 is painful or useless

The main sequence is the diagonal from (A=0, I=1) to (A=1, I=0): - Zone of pain: stable (low I) but concrete (low A) — hard to change, many depend on it - Zone of uselessness: abstract (high A) but unused (high I) — defines contracts but nothing uses them

Plugin Seams: Extractable Architectures

A seam is a contract plus the modules that implement it. It answers: "What can be lifted out whole?"

How Seams Are Found

  1. A module is a contract if it declares abstract types (Protocol, interface, trait, ABC)
  2. A module implements the contract if it declares a type deriving from it
  3. A module consumes the contract if it merely depends on it
  4. The unit = contract + implementers; escapes = dependencies outside the unit
  5. Cohesion = |unit| / (|unit| + |escapes|) predicts extraction safety

Verdicts

  • clean (cohesion ≥ 0.9) — Extract with confidence
  • extractable (≥ 0.6) — Safe to extract; assess cost of escaping dependencies
  • entangled (< 0.6) — Declared plug point, but implementations are entangled
  • declared-unused — Zero implementers; not actually a plug point

Why Seams Matter

The same structural property makes something both extractable and pluggable: - Extracting: Take the contract + implementers, leave consumers behind - Recomposing: Consumers stay outside, depending on the contract — the plug point

Evidence from real projects:

  • Drydock's langs/base.py scores 0.88 with 1 escape. This component was successfully extracted into a standalone package; that one escape was exactly what came along.
  • Drydock's extract/rewriters/base.py scores 1.0 with zero escapes — a perfect extraction candidate.
  • A real Go project has plugin/plugin.go with 27 dependents but 45 escaping dependencies. The verdict is entangled — a plug point exists on paper, but the implementations are not decoupled.

Finding Seams

# List all seams, ranked by cohesion
python3 -m drydock.cli extract <project> --list-seams

# Show ready-to-run extraction commands
python3 -m drydock.cli extract <project> --list-seams --markdown

See Plugin Seams for deep conceptual details and Extract for how to materialize them.

Deferred Imports, Not Cycles

A critical distinction for accuracy:

Load-time import (module-level import statement) — creates a true circular dependency.

Deferred import (inside a function or if TYPE_CHECKING: block) — creates no load-time dependency; a cycle through one is a deliberate inversion, the standard way a registry imports its own plugins.

Example: Drydock's core/registry.py imports its own plugins inside a function. Naive analysis reports a cycle; this analyzer reports 0 load-time cycles and 14 deferred-only ones, correctly identifying the inversion.

Example Output

Markdown (Human-Readable)

# Architecture: drydock

46 modules · 110 load-time dependencies · 16 deferred · 8 levels

## Principles

### Acyclic Dependencies — holds

No circular dependencies at load time. 14 cycle(s) exist only through deferred imports (inside functions or TYPE_CHECKING blocks), which is the standard way to invert a registry dependency deliberately.

<details><summary>evidence</summary>

- `load_time_cycles`: 0
- `deferred_only_cycles`: 14
- `load_time_edges`: 110
- `deferred_edges`: 16

</details>

### Stable Dependencies Principle — holds

109 of 110 load-time dependencies point toward equally or more stable modules.

<details><summary>evidence</summary>

- `edges`: 110
- `violating_edges`: 1
- `violation_ratio`: 0.009

</details>

- from=extract/__init__.py, to=extract/planner.py, from_instability=0.667, to_instability=0.75

## Most depended upon

| Module | Ca | Ce | I | A | D |
|---|---|---|---|---|---|
| `core/models.py` | 19 | 0 | 0.0 | 0.0 | 1.0 |
| `core/registry.py` | 18 | 6 | 0.25 | 0.25 | 0.5 |
| `core/output.py` | 12 | 0 | 0.0 | 0.0 | 1.0 |

JSON (Machine-Readable)

{
  "meta": {
    "project": "drydock",
    "root": "/path/to/drydock",
    "files_walked": 46,
    "truncated": false
  },
  "graph": {
    "modules": 46,
    "edges": 126,
    "load_time_edges": 110,
    "deferred_edges": 16,
    "cycles": 0,
    "layers": 8
  },
  "principles": [
    {
      "id": "acyclic-dependencies",
      "name": "Acyclic Dependencies",
      "verdict": "holds",
      "summary": "No circular dependencies at load time.",
      "derived": "deterministic",
      "evidence": {
        "load_time_cycles": 0,
        "deferred_only_cycles": 14,
        "load_time_edges": 110,
        "deferred_edges": 16
      },
      "violations": []
    }
  ],
  "verdicts": {
    "holds": 5,
    "violated": 0,
    "partial": 1,
    "not_applicable": 1
  },
  "metrics": {
    "per_module": [...],
    "most_depended_upon": [...],
    "furthest_from_main_sequence": [...]
  }
}

Verdict Values

Each principle returns one of these verdicts:

  • holds — the principle is satisfied
  • violated — the principle is clearly not satisfied
  • partial — the principle is satisfied but with notable exceptions
  • not_applicable — the codebase gives no signal (too small, no types, etc.)

Options

Flag Type Default Description
project_path positional required Path to project root
--output, -o string stdout Output file path
--markdown, -m flag false Output as Markdown (default: JSON)
--compact, -c flag false Compact JSON, no indentation
--no-tests flag false Exclude test files from analysis
--max-files int unlimited Stop after N files (partial results flagged)
--skip-dir string repeatable Additional directory name to skip
--follow-symlinks flag false Follow symlinked directories
--min-evidence int 0 Omit findings whose evidence covers fewer than N modules
--include-metrics flag true Include the per-module metrics table
--top int 15 How many modules to list in ranked tables

Supported Languages

All metrics work across all supported languages. Type detection (abstractness) uses language-native markers only:

  • Python: Protocol, ABC, @abstractmethod
  • Go: interface
  • Rust: trait
  • Java/C#: interface
  • TypeScript: interface

Classes without a language-native marker (e.g., a class named AbstractFoo) count as concrete, deliberately. This prevents false positives from naming conventions.

Limits & Caveats

Type Counting is Approximate for Non-Python

Python uses AST for precise symbol extraction. Other languages use regex, so class/type counts are approximate. This doesn't affect cycle detection, dependency flow, or layering — only the abstractness metric.

Abstractness Requires Type Definitions

Modules with no classes or types report abstractness = 0. The principle only evaluates modules that define types (>0 in the evidence).

Verdict Thresholds are Conventions

Thresholds are sensible defaults (e.g., Stable Dependencies tolerates ≤5% violating edges), but they are not gospel. The raw numbers are always in the evidence so a reader can apply their own threshold.

The Analyzer Describes Structure, Not Intent

The analyzer cannot tell you why a design is the way it is. It can tell you that 3 modules are heavily depended upon but concrete (zone of pain); you must decide whether that's justified by their role or an opportunity to extract interfaces.

Use Cases

  • Learning from an unfamiliar codebase: "Is this layered? Where are the hubs? Are they abstract?"
  • Comparing codebase designs: Run on several projects and compare their principle verdicts
  • Before/after refactoring: Check whether a refactoring improved the architecture (fewer cycles, better abstraction)
  • Pre-extraction review: Verify stability and coupling before extracting a component
  • CI/CD gating: Fail the build if principles change unexpectedly

Cross-Language Examples

Python (Drydock)

46 modules, 110 load-time edges, 0 cycles, 8 levels
Acyclic: holds
Layering: holds (8 levels)
Stable Dependencies: holds (109/110 edges correct)
Stable Abstractions: holds (5 in zone of pain, not > 25%)
Dependency Inversion: holds (1 of 1 hubs define contracts)
No God Modules: holds (no bottlenecks)
Tests Target Public: partial (7/22 test files reach into internals)

Go (Fabric, 274 modules)

274 modules, 1387 load-time edges, 0 cycles, 9 levels
Acyclic: holds
Layering: holds (9 levels, 87/96 packages aligned)
Stable Dependencies: holds (99.2% of edges correct)
Dependency Inversion: holds (4/5 hubs define interfaces)
No God Modules: holds (threshold 68.5, no module exceeds it)

Further Reading