Extending Drydock¶
The headline feature of Drydock 0.2 is the plugin system. Extend Drydock by implementing the Analyzer and LanguageParser protocols — the same contracts the built-in tools use.
Building Tools on the Catalog¶
The Catalog class is the shared API for building TUIs, GUIs, and automation layers on top of Drydock. Rather than accessing the SQLite database directly, queries return plain dataclasses that are:
- Serializable — Convert to JSON for transmission over MCP, REST, or sockets
- Detachable — Results can be held in memory, passed between processes, or stored without a database connection
- Testable — No database setup needed for unit tests
A GUI, TUI, or web service can instantiate a Catalog, perform queries, and hand results to UI rendering logic without coupling to SQLite or worrying about connection management.
Writing a Custom Analyzer¶
An analyzer declares its options once, as data. The CLI turns that declaration into argparse flags and the MCP server turns it into a typed tool signature. Nothing is written twice, so they cannot drift.
Minimal Example¶
from drydock.core.registry import Analyzer, AnalysisContext, Option
class MyAnalyzer:
"""Analyze something custom about the codebase."""
name = "my_analyzer"
summary = "One-line summary shown in drydock list"
description = """Full description for AI agents. Explains what this analyzer
does and when to use it."""
options = (
Option(
name="min_size",
type=int,
default=100,
help="Minimum file size in bytes to report",
),
Option(
name="include_config",
type=bool,
default=False,
help="Include configuration files in analysis",
),
)
def run(self, ctx: AnalysisContext, **opts) -> dict:
"""Produce the analysis.
Args:
ctx: Shared analysis context (walks tree once, caches symbols)
**opts: User-supplied options (min_size=100, include_config=False, etc.)
Returns:
Dict containing analysis results. Will be serialized to JSON by default.
"""
min_size = opts.get("min_size", 100)
include_config = opts.get("include_config", False)
results = []
for source_file in ctx.source_files():
if source_file.size >= min_size:
results.append({
"path": str(source_file.id),
"size": source_file.size,
"lang": source_file.lang,
"lines": source_file.line_count,
})
return {
"meta": ctx.meta(),
"files_analyzed": len(results),
"files": results,
}
def to_markdown(self, result: dict) -> str:
"""Render results as human-readable markdown.
Args:
result: Dict returned by run()
Returns:
Markdown string. Every section of the JSON should appear.
"""
lines = [
f"# Custom Analyzer Results\n",
f"Files analyzed: {result['files_analyzed']}\n",
"\n## Files\n",
]
for file in result["files"]:
lines.append(
f"- `{file['path']}` ({file['size']} bytes, "
f"{file['lines']} lines, {file['lang']})\n"
)
return "".join(lines)
Register the Analyzer¶
In your pyproject.toml:
Install your package:
Now it's available:
drydock list
# ...
# my_analyzer One-line summary shown in drydock list
drydock my_analyzer /path/to/project
drydock my_analyzer /path/to/project --min-size 500
drydock my_analyzer /path/to/project --include-config --markdown
It also appears in the MCP server:
# From Claude Code or any MCP-compatible assistant:
# `drydock_my_analyzer` is now available as a native tool call
Writing a Custom Architectural Principle¶
An architectural principle evaluates one aspect of the dependency graph. It answers a yes/no/partial question about a codebase by examining the import graph and symbol table.
A principle is deterministic — given the same import graph, it always returns the same verdict with the same evidence. This is key: a verdict that cannot be verified is worse than no verdict.
The Finding Contract¶
Every principle must return a Finding with mandatory evidence. A Finding without evidence is not a finding.
@dataclass(frozen=True, slots=True)
class Finding:
id: str # e.g. "my-principle"
name: str # e.g. "My Principle"
verdict: Verdict # "holds" | "violated" | "partial" | "not_applicable"
summary: str # One sentence stating the fact
evidence: dict[str, Any] # The numbers and module lists it was computed from
violations: tuple[dict[str, Any], ...] = () # Concrete counterexamples
derived: Literal["deterministic", "heuristic"] = "deterministic"
Minimal Example¶
from dataclasses import dataclass
from drydock.arch.graph import DependencyGraph
from drydock.arch.principle import Finding, Principle, principles
@dataclass
class MyPrinciple:
"""Measure something about the dependency graph."""
id = "my-principle"
name = "My Principle"
question = "Do dependencies flow toward the center?"
def evaluate(self, graph: DependencyGraph) -> Finding | None:
"""Evaluate the principle.
Args:
graph: The resolved dependency graph
Returns:
A Finding with verdict and evidence, or None if not applicable
"""
if not graph.edges:
return None
# Compute something about the graph
violations = []
for module, metrics in graph.metrics.items():
if metrics.ce > 10: # Example: modules with too many dependencies
violations.append({
"module": module,
"outbound_edges": metrics.ce
})
return Finding(
id=self.id,
name=self.name,
verdict="holds" if not violations else "violated",
summary=f"{len(violations)} modules have excessive outbound coupling.",
evidence={
"modules_total": len(graph.modules),
"violation_count": len(violations),
"threshold": 10,
},
violations=tuple(violations[:10]),
)
# Register the principle
principles.register(MyPrinciple())
Register the Principle¶
In your pyproject.toml:
Install:
Now it appears in architecture analysis:
drydock architecture /path/to/project --markdown
# ... includes your principle in the output
drydock architecture /path/to/project | jq '.principles[] | select(.id == "my-principle")'
Best Practices¶
-
Evidence is Mandatory — Every Finding must carry the numbers and module lists it was computed from. If a reader cannot verify the verdict, the principle is not trustworthy.
-
One Question Only — A principle answers one architectural question. Keep it focused.
-
Deterministic — Same graph, same verdict, every time. No randomness, no thresholds that change.
-
Handle Edge Cases — Return
not_applicableif the codebase gives no signal (too small, no types, no imports, etc.). -
Name Modules Concretely — Evidence should include module names, not abstract counts. A reader learning from your principle should be able to inspect those modules.
Accessing the Graph¶
The DependencyGraph gives you:
graph.modules # dict[FileId, SourceFile] — all files
graph.edges # list[Edge] — all import edges
graph.load_time_edges() # list[Edge] — only non-deferred imports
graph.cycles() # list[list[FileId]] — elementary cycles (load-time by default)
graph.layers() # dict[int, list[FileId]] — topological levels
graph.metrics # dict[FileId, ModuleMetrics] — per-module metrics
graph.package_metrics() # dict[str, ModuleMetrics] — aggregated to packages
Each ModuleMetrics includes:
- ca: how many modules depend on this one (afferent coupling)
- ce: how many modules this depends on (efferent coupling)
- instability: Ce / (Ca + Ce), where 0 = stable, 1 = unstable
- abstractness: (abstract types) / (all types)
- distance: |A + I - 1|, distance from the main sequence
- depth: topological depth over load-time edges
Testing Your Principle¶
import pytest
from drydock.arch.graph import DependencyGraph, Edge, ModuleMetrics
from mypackage.arch import MyPrinciple
def test_my_principle_holds():
# Build a minimal graph
graph = DependencyGraph()
graph.modules = {
"a.py": None, # SourceFile, can be None for testing
"b.py": None,
}
graph.edges = [Edge("a.py", "b.py")]
graph.metrics = {
"a.py": ModuleMetrics("a.py", ca=0, ce=1, abstract_types=0,
concrete_types=1, depth=1),
"b.py": ModuleMetrics("b.py", ca=1, ce=0, abstract_types=1,
concrete_types=0, depth=0),
}
principle = MyPrinciple()
finding = principle.evaluate(graph)
assert finding is not None
assert finding.verdict == "holds"
assert "violation_count" in finding.evidence
assert finding.evidence["violation_count"] == 0
def test_my_principle_violated():
# Build a graph where the principle is violated
graph = DependencyGraph()
graph.modules = {"a.py": None}
graph.edges = []
graph.metrics = {
"a.py": ModuleMetrics("a.py", ca=0, ce=15, abstract_types=0,
concrete_types=1, depth=0),
}
principle = MyPrinciple()
finding = principle.evaluate(graph)
assert finding.verdict == "violated"
assert finding.evidence["violation_count"] == 1
assert finding.violations[0]["module"] == "a.py"
Writing a Custom Language Backend¶
A language backend answers three questions about one file:
symbols(...)— What is defined here?imports(...)— What does it reference?interface(...)— What does it expose?
It does not decide where imports point — that is ModuleIndex's job. Keeping the split means a new language integrates with all analyzers at once: register a parser and codemap, boundaries, dependencies, and interfaces all pick it up.
Minimal Example¶
from drydock.langs.base import BaseParser
from drydock.core.models import Symbols, ImportRef, Interface
class MyLanguageParser(BaseParser):
name = "mylang"
extensions = (".ml", ".mli") # Example: OCaml
fidelity = "regex" # or "ast" if using a tree-sitter parser
def symbols(self, source: str, path: str) -> Symbols:
"""Extract all definitions: classes, functions, constants, etc.
Args:
source: File contents as string
path: File path (for error messages)
Returns:
Symbols object containing all definitions and imports
"""
# Example: parse with regex
import re
imports = []
functions = []
classes = []
# Find imports: open Foo
for match in re.finditer(r'open\s+(\w+)', source):
imports.append(ImportRef(
name=match.group(1),
lineno=source[:match.start()].count('\n') + 1,
))
# Find functions: let foo = ...
for match in re.finditer(r'let\s+(\w+)\s*=', source):
functions.append({
"name": match.group(1),
"lineno": source[:match.start()].count('\n') + 1,
})
return Symbols(
imports=tuple(imports),
functions=functions,
classes=classes,
constants=[],
type_defs=[],
exports=[],
)
def imports(self, source: str, path: str) -> tuple[ImportRef, ...]:
"""Extract only imports (names of things referenced).
Default implementation returns self.symbols().imports.
Override if your language has a different way to extract imports.
"""
return self.symbols(source, path).imports
def interface(self, source: str, path: str) -> Interface:
"""Extract public API surface of this file.
Default implementation returns public identifiers by naming convention.
Override if your language has explicit export syntax.
"""
sym = self.symbols(source, path)
# By default, anything not prefixed with _ is public
return Interface.from_symbols(sym)
Register the Language Backend¶
In your pyproject.toml:
Install:
Now it's available in all analyzers:
drydock codemap /path/to/mylang/project
# Automatically detects .ml and .mli files
drydock languages
# Shows your language with fidelity level
The AnalysisContext¶
Every analyzer receives an AnalysisContext — a shared, cached view of the project:
def run(self, ctx: AnalysisContext, **opts) -> dict:
# Walk the tree once, files are cached
walk = ctx.walk # WalkResult with all files
# Build module index once (resolves imports)
index = ctx.index # ModuleIndex
# Get parsed symbols for a file (cached)
symbols = ctx.symbols(source_file)
# Get public interface (cached)
interface = ctx.interface(source_file)
# Get file imports (cached, via symbols)
imports = ctx.imports(source_file)
# Get just source files (not build artifacts)
sources = ctx.source_files()
# Metadata for your results
meta = ctx.meta() # {"project": "...", "files_walked": N, ...}
# Report progress to the user
ctx.reporter.step("doing something")
ctx.reporter.sub("substep")
ctx.reporter.warn("something might be wrong")
Best Practices¶
-
Options are data — Declare them once in
options = (...). The CLI and MCP server read this, so they stay in sync. -
Render complete markdown — Your
to_markdown()must include every section of the JSON. Pre-0.2 shipped an interfaces section that was silently dropped when rendering markdown. -
Non-destructive — Never modify the project. Your analyzer is a reporter, not an editor.
-
Cache-aware —
ctx.symbols(),ctx.interface(), andctx.source()are cached per-run. If you call them multiple times on the same file, you're reading the cache. -
Avoid side effects — Keep your analyzer pure. Deterministic output makes it reproducible for CI/CD.
-
Compose with others — Your analyzer runs alongside others. Don't assume you're the only one touching the tree.
Testing Your Analyzer¶
import pytest
from pathlib import Path
from drydock.core.registry import AnalysisContext
from mypackage.analyzers import MyAnalyzer
def test_my_analyzer(tmp_path: Path):
# Create a minimal test project
test_file = tmp_path / "test.py"
test_file.write_text("def foo(): pass")
ctx = AnalysisContext(tmp_path)
analyzer = MyAnalyzer()
result = analyzer.run(ctx, min_size=0)
assert result["files_analyzed"] == 1
assert len(result["files"]) == 1
assert result["files"][0]["path"].endswith("test.py")
def test_markdown_rendering():
analyzer = MyAnalyzer()
result = {
"meta": {"project": "test", "files_walked": 1},
"files_analyzed": 1,
"files": [
{"path": "foo.py", "size": 100, "lang": "python", "lines": 5}
],
}
markdown = analyzer.to_markdown(result)
assert "foo.py" in markdown
assert "100 bytes" in markdown
Writing a Custom Import Rewriter¶
When extracting components, imports must be repointed so the extracted copy resolves correctly. An import rewriter is a pure function that takes source text and a module mapping, and returns rewritten source.
Rewriters are language-specific and designed to be conservative: when unsure, leave the import alone. This keeps extraction non-destructive.
Minimal Example¶
from drydock.extract.rewriters.base import BaseRewriter
class CustomLangRewriter(BaseRewriter):
language = "customlang"
def rewrite(
self, source: str, path: str, module_map: dict[str, str]
) -> tuple[str, int]:
"""Rewrite imports in one file.
Args:
source: File contents as string
path: File path (for logging/context)
module_map: Mapping of old_module -> new_module
(e.g., {"drydock.langs": "langpack.langs"})
Returns:
Tuple of (rewritten_source, number_of_rewrites_made).
If no rewrites are made, return (source, 0).
"""
import re
# Example: rewrite "import drydock.langs" to "import langpack.langs"
rewritten = source
count = 0
for old_module, new_module in module_map.items():
# Match "import old_module" or "from old_module"
pattern = rf'\b{re.escape(old_module)}\b'
new_source = re.sub(
pattern,
new_module,
rewritten,
)
count += len(re.findall(pattern, rewritten))
rewritten = new_source
return (rewritten, count)
Register the Rewriter¶
In your pyproject.toml:
Install and test:
pip install -e .
# The rewriter is now used automatically during extraction
python3 -m drydock.cli extract /path/to/project --dir component_path -o output
Best Practices¶
-
Conservative — Prefer leaving imports alone over making a wrong rewrite. Broken imports in the extracted copy are better than silent corruption.
-
Whole segments only — Match dotted names segment-by-segment (
pkg.subnotpkgsub), so rewrites don't accidentally match prefixes. -
Longest-prefix-first — If both
pkgandpkg.submap, preferpkg.sub. TheBaseRewriter.longest_prefix_match()helper does this. -
Return accurate counts — The count is reported in the extraction manifest and plan, so it must be correct.
-
Deterministic — Same source + same mapping = same output, every time. No side effects or state.
Testing Your Rewriter¶
def test_customlang_rewriter():
rewriter = CustomLangRewriter()
source = """
import drydock.langs
from drydock.core import models
"""
module_map = {
"drydock.langs": "langpack.langs",
"drydock.core": "langpack.core",
}
rewritten, count = rewriter.rewrite(source, "test.customlang", module_map)
assert "langpack.langs" in rewritten
assert "langpack.core" in rewritten
assert count == 2
Examples¶
- Built-in analyzers — Reference implementations
- Built-in language backends — Python AST, regex examples
- Built-in import rewriters — Python, JavaScript, Go, Rust