Skip to content

Extending Drydock

The headline feature of Drydock 0.2 is the plugin system. Extend Drydock by implementing the Analyzer and LanguageParser protocols — the same contracts the built-in tools use.

Building Tools on the Catalog

The Catalog class is the shared API for building TUIs, GUIs, and automation layers on top of Drydock. Rather than accessing the SQLite database directly, queries return plain dataclasses that are:

  • Serializable — Convert to JSON for transmission over MCP, REST, or sockets
  • Detachable — Results can be held in memory, passed between processes, or stored without a database connection
  • Testable — No database setup needed for unit tests

A GUI, TUI, or web service can instantiate a Catalog, perform queries, and hand results to UI rendering logic without coupling to SQLite or worrying about connection management.

Writing a Custom Analyzer

An analyzer declares its options once, as data. The CLI turns that declaration into argparse flags and the MCP server turns it into a typed tool signature. Nothing is written twice, so they cannot drift.

Minimal Example

from drydock.core.registry import Analyzer, AnalysisContext, Option

class MyAnalyzer:
    """Analyze something custom about the codebase."""

    name = "my_analyzer"
    summary = "One-line summary shown in drydock list"
    description = """Full description for AI agents. Explains what this analyzer
    does and when to use it."""

    options = (
        Option(
            name="min_size",
            type=int,
            default=100,
            help="Minimum file size in bytes to report",
        ),
        Option(
            name="include_config",
            type=bool,
            default=False,
            help="Include configuration files in analysis",
        ),
    )

    def run(self, ctx: AnalysisContext, **opts) -> dict:
        """Produce the analysis.

        Args:
            ctx: Shared analysis context (walks tree once, caches symbols)
            **opts: User-supplied options (min_size=100, include_config=False, etc.)

        Returns:
            Dict containing analysis results. Will be serialized to JSON by default.
        """
        min_size = opts.get("min_size", 100)
        include_config = opts.get("include_config", False)

        results = []
        for source_file in ctx.source_files():
            if source_file.size >= min_size:
                results.append({
                    "path": str(source_file.id),
                    "size": source_file.size,
                    "lang": source_file.lang,
                    "lines": source_file.line_count,
                })

        return {
            "meta": ctx.meta(),
            "files_analyzed": len(results),
            "files": results,
        }

    def to_markdown(self, result: dict) -> str:
        """Render results as human-readable markdown.

        Args:
            result: Dict returned by run()

        Returns:
            Markdown string. Every section of the JSON should appear.
        """
        lines = [
            f"# Custom Analyzer Results\n",
            f"Files analyzed: {result['files_analyzed']}\n",
            "\n## Files\n",
        ]
        for file in result["files"]:
            lines.append(
                f"- `{file['path']}` ({file['size']} bytes, "
                f"{file['lines']} lines, {file['lang']})\n"
            )
        return "".join(lines)

Register the Analyzer

In your pyproject.toml:

[project.entry-points."drydock.analyzers"]
my_analyzer = "mypackage.analyzers:MyAnalyzer"

Install your package:

pip install -e .

Now it's available:

drydock list
# ...
# my_analyzer  One-line summary shown in drydock list

drydock my_analyzer /path/to/project
drydock my_analyzer /path/to/project --min-size 500
drydock my_analyzer /path/to/project --include-config --markdown

It also appears in the MCP server:

# From Claude Code or any MCP-compatible assistant:
# `drydock_my_analyzer` is now available as a native tool call

Writing a Custom Architectural Principle

An architectural principle evaluates one aspect of the dependency graph. It answers a yes/no/partial question about a codebase by examining the import graph and symbol table.

A principle is deterministic — given the same import graph, it always returns the same verdict with the same evidence. This is key: a verdict that cannot be verified is worse than no verdict.

The Finding Contract

Every principle must return a Finding with mandatory evidence. A Finding without evidence is not a finding.

@dataclass(frozen=True, slots=True)
class Finding:
    id: str                  # e.g. "my-principle"
    name: str                # e.g. "My Principle"
    verdict: Verdict         # "holds" | "violated" | "partial" | "not_applicable"
    summary: str             # One sentence stating the fact
    evidence: dict[str, Any] # The numbers and module lists it was computed from
    violations: tuple[dict[str, Any], ...] = ()  # Concrete counterexamples
    derived: Literal["deterministic", "heuristic"] = "deterministic"

Minimal Example

from dataclasses import dataclass
from drydock.arch.graph import DependencyGraph
from drydock.arch.principle import Finding, Principle, principles

@dataclass
class MyPrinciple:
    """Measure something about the dependency graph."""

    id = "my-principle"
    name = "My Principle"
    question = "Do dependencies flow toward the center?"

    def evaluate(self, graph: DependencyGraph) -> Finding | None:
        """Evaluate the principle.

        Args:
            graph: The resolved dependency graph

        Returns:
            A Finding with verdict and evidence, or None if not applicable
        """
        if not graph.edges:
            return None

        # Compute something about the graph
        violations = []
        for module, metrics in graph.metrics.items():
            if metrics.ce > 10:  # Example: modules with too many dependencies
                violations.append({
                    "module": module,
                    "outbound_edges": metrics.ce
                })

        return Finding(
            id=self.id,
            name=self.name,
            verdict="holds" if not violations else "violated",
            summary=f"{len(violations)} modules have excessive outbound coupling.",
            evidence={
                "modules_total": len(graph.modules),
                "violation_count": len(violations),
                "threshold": 10,
            },
            violations=tuple(violations[:10]),
        )

# Register the principle
principles.register(MyPrinciple())

Register the Principle

In your pyproject.toml:

[project.entry-points."drydock.principles"]
my_principle = "mypackage.arch:MyPrinciple"

Install:

pip install -e .

Now it appears in architecture analysis:

drydock architecture /path/to/project --markdown
# ... includes your principle in the output

drydock architecture /path/to/project | jq '.principles[] | select(.id == "my-principle")'

Best Practices

  1. Evidence is Mandatory — Every Finding must carry the numbers and module lists it was computed from. If a reader cannot verify the verdict, the principle is not trustworthy.

  2. One Question Only — A principle answers one architectural question. Keep it focused.

  3. Deterministic — Same graph, same verdict, every time. No randomness, no thresholds that change.

  4. Handle Edge Cases — Return not_applicable if the codebase gives no signal (too small, no types, no imports, etc.).

  5. Name Modules Concretely — Evidence should include module names, not abstract counts. A reader learning from your principle should be able to inspect those modules.

Accessing the Graph

The DependencyGraph gives you:

graph.modules            # dict[FileId, SourceFile] — all files
graph.edges              # list[Edge] — all import edges
graph.load_time_edges()  # list[Edge] — only non-deferred imports
graph.cycles()           # list[list[FileId]] — elementary cycles (load-time by default)
graph.layers()           # dict[int, list[FileId]] — topological levels
graph.metrics            # dict[FileId, ModuleMetrics] — per-module metrics
graph.package_metrics()  # dict[str, ModuleMetrics] — aggregated to packages

Each ModuleMetrics includes: - ca: how many modules depend on this one (afferent coupling) - ce: how many modules this depends on (efferent coupling) - instability: Ce / (Ca + Ce), where 0 = stable, 1 = unstable - abstractness: (abstract types) / (all types) - distance: |A + I - 1|, distance from the main sequence - depth: topological depth over load-time edges

Testing Your Principle

import pytest
from drydock.arch.graph import DependencyGraph, Edge, ModuleMetrics
from mypackage.arch import MyPrinciple

def test_my_principle_holds():
    # Build a minimal graph
    graph = DependencyGraph()
    graph.modules = {
        "a.py": None,  # SourceFile, can be None for testing
        "b.py": None,
    }
    graph.edges = [Edge("a.py", "b.py")]
    graph.metrics = {
        "a.py": ModuleMetrics("a.py", ca=0, ce=1, abstract_types=0, 
                              concrete_types=1, depth=1),
        "b.py": ModuleMetrics("b.py", ca=1, ce=0, abstract_types=1, 
                              concrete_types=0, depth=0),
    }

    principle = MyPrinciple()
    finding = principle.evaluate(graph)

    assert finding is not None
    assert finding.verdict == "holds"
    assert "violation_count" in finding.evidence
    assert finding.evidence["violation_count"] == 0

def test_my_principle_violated():
    # Build a graph where the principle is violated
    graph = DependencyGraph()
    graph.modules = {"a.py": None}
    graph.edges = []
    graph.metrics = {
        "a.py": ModuleMetrics("a.py", ca=0, ce=15, abstract_types=0, 
                              concrete_types=1, depth=0),
    }

    principle = MyPrinciple()
    finding = principle.evaluate(graph)

    assert finding.verdict == "violated"
    assert finding.evidence["violation_count"] == 1
    assert finding.violations[0]["module"] == "a.py"

Writing a Custom Language Backend

A language backend answers three questions about one file:

  • symbols(...) — What is defined here?
  • imports(...) — What does it reference?
  • interface(...) — What does it expose?

It does not decide where imports point — that is ModuleIndex's job. Keeping the split means a new language integrates with all analyzers at once: register a parser and codemap, boundaries, dependencies, and interfaces all pick it up.

Minimal Example

from drydock.langs.base import BaseParser
from drydock.core.models import Symbols, ImportRef, Interface

class MyLanguageParser(BaseParser):
    name = "mylang"
    extensions = (".ml", ".mli")  # Example: OCaml
    fidelity = "regex"  # or "ast" if using a tree-sitter parser

    def symbols(self, source: str, path: str) -> Symbols:
        """Extract all definitions: classes, functions, constants, etc.

        Args:
            source: File contents as string
            path: File path (for error messages)

        Returns:
            Symbols object containing all definitions and imports
        """
        # Example: parse with regex
        import re
        imports = []
        functions = []
        classes = []

        # Find imports: open Foo
        for match in re.finditer(r'open\s+(\w+)', source):
            imports.append(ImportRef(
                name=match.group(1),
                lineno=source[:match.start()].count('\n') + 1,
            ))

        # Find functions: let foo = ...
        for match in re.finditer(r'let\s+(\w+)\s*=', source):
            functions.append({
                "name": match.group(1),
                "lineno": source[:match.start()].count('\n') + 1,
            })

        return Symbols(
            imports=tuple(imports),
            functions=functions,
            classes=classes,
            constants=[],
            type_defs=[],
            exports=[],
        )

    def imports(self, source: str, path: str) -> tuple[ImportRef, ...]:
        """Extract only imports (names of things referenced).

        Default implementation returns self.symbols().imports.
        Override if your language has a different way to extract imports.
        """
        return self.symbols(source, path).imports

    def interface(self, source: str, path: str) -> Interface:
        """Extract public API surface of this file.

        Default implementation returns public identifiers by naming convention.
        Override if your language has explicit export syntax.
        """
        sym = self.symbols(source, path)
        # By default, anything not prefixed with _ is public
        return Interface.from_symbols(sym)

Register the Language Backend

In your pyproject.toml:

[project.entry-points."drydock.languages"]
mylang = "mypackage.langs:MyLanguageParser"

Install:

pip install -e .

Now it's available in all analyzers:

drydock codemap /path/to/mylang/project
# Automatically detects .ml and .mli files

drydock languages
# Shows your language with fidelity level

The AnalysisContext

Every analyzer receives an AnalysisContext — a shared, cached view of the project:

def run(self, ctx: AnalysisContext, **opts) -> dict:
    # Walk the tree once, files are cached
    walk = ctx.walk  # WalkResult with all files

    # Build module index once (resolves imports)
    index = ctx.index  # ModuleIndex

    # Get parsed symbols for a file (cached)
    symbols = ctx.symbols(source_file)

    # Get public interface (cached)
    interface = ctx.interface(source_file)

    # Get file imports (cached, via symbols)
    imports = ctx.imports(source_file)

    # Get just source files (not build artifacts)
    sources = ctx.source_files()

    # Metadata for your results
    meta = ctx.meta()  # {"project": "...", "files_walked": N, ...}

    # Report progress to the user
    ctx.reporter.step("doing something")
    ctx.reporter.sub("substep")
    ctx.reporter.warn("something might be wrong")

Best Practices

  1. Options are data — Declare them once in options = (...). The CLI and MCP server read this, so they stay in sync.

  2. Render complete markdown — Your to_markdown() must include every section of the JSON. Pre-0.2 shipped an interfaces section that was silently dropped when rendering markdown.

  3. Non-destructive — Never modify the project. Your analyzer is a reporter, not an editor.

  4. Cache-aware — ctx.symbols(), ctx.interface(), and ctx.source() are cached per-run. If you call them multiple times on the same file, you're reading the cache.

  5. Avoid side effects — Keep your analyzer pure. Deterministic output makes it reproducible for CI/CD.

  6. Compose with others — Your analyzer runs alongside others. Don't assume you're the only one touching the tree.

Testing Your Analyzer

import pytest
from pathlib import Path
from drydock.core.registry import AnalysisContext
from mypackage.analyzers import MyAnalyzer

def test_my_analyzer(tmp_path: Path):
    # Create a minimal test project
    test_file = tmp_path / "test.py"
    test_file.write_text("def foo(): pass")

    ctx = AnalysisContext(tmp_path)
    analyzer = MyAnalyzer()
    result = analyzer.run(ctx, min_size=0)

    assert result["files_analyzed"] == 1
    assert len(result["files"]) == 1
    assert result["files"][0]["path"].endswith("test.py")

def test_markdown_rendering():
    analyzer = MyAnalyzer()
    result = {
        "meta": {"project": "test", "files_walked": 1},
        "files_analyzed": 1,
        "files": [
            {"path": "foo.py", "size": 100, "lang": "python", "lines": 5}
        ],
    }
    markdown = analyzer.to_markdown(result)
    assert "foo.py" in markdown
    assert "100 bytes" in markdown

Writing a Custom Import Rewriter

When extracting components, imports must be repointed so the extracted copy resolves correctly. An import rewriter is a pure function that takes source text and a module mapping, and returns rewritten source.

Rewriters are language-specific and designed to be conservative: when unsure, leave the import alone. This keeps extraction non-destructive.

Minimal Example

from drydock.extract.rewriters.base import BaseRewriter

class CustomLangRewriter(BaseRewriter):
    language = "customlang"

    def rewrite(
        self, source: str, path: str, module_map: dict[str, str]
    ) -> tuple[str, int]:
        """Rewrite imports in one file.

        Args:
            source: File contents as string
            path: File path (for logging/context)
            module_map: Mapping of old_module -> new_module
                       (e.g., {"drydock.langs": "langpack.langs"})

        Returns:
            Tuple of (rewritten_source, number_of_rewrites_made).
            If no rewrites are made, return (source, 0).
        """
        import re

        # Example: rewrite "import drydock.langs" to "import langpack.langs"
        rewritten = source
        count = 0

        for old_module, new_module in module_map.items():
            # Match "import old_module" or "from old_module"
            pattern = rf'\b{re.escape(old_module)}\b'
            new_source = re.sub(
                pattern,
                new_module,
                rewritten,
            )
            count += len(re.findall(pattern, rewritten))
            rewritten = new_source

        return (rewritten, count)

Register the Rewriter

In your pyproject.toml:

[project.entry-points."drydock.rewriters"]
customlang = "mypackage.rewriters:CustomLangRewriter"

Install and test:

pip install -e .

# The rewriter is now used automatically during extraction
python3 -m drydock.cli extract /path/to/project --dir component_path -o output

Best Practices

  1. Conservative — Prefer leaving imports alone over making a wrong rewrite. Broken imports in the extracted copy are better than silent corruption.

  2. Whole segments only — Match dotted names segment-by-segment (pkg.sub not pkgsub), so rewrites don't accidentally match prefixes.

  3. Longest-prefix-first — If both pkg and pkg.sub map, prefer pkg.sub. The BaseRewriter.longest_prefix_match() helper does this.

  4. Return accurate counts — The count is reported in the extraction manifest and plan, so it must be correct.

  5. Deterministic — Same source + same mapping = same output, every time. No side effects or state.

Testing Your Rewriter

def test_customlang_rewriter():
    rewriter = CustomLangRewriter()

    source = """
    import drydock.langs
    from drydock.core import models
    """

    module_map = {
        "drydock.langs": "langpack.langs",
        "drydock.core": "langpack.core",
    }

    rewritten, count = rewriter.rewrite(source, "test.customlang", module_map)

    assert "langpack.langs" in rewritten
    assert "langpack.core" in rewritten
    assert count == 2

Examples