Skip to content

Catalog: Your Accumulated Component Library

The catalog is a persistent library that accumulates — not a cache. It records repositories at specific commits, architectural facts with evidence, extracted components, and the public symbols of every module. Once ingested, a project's structure and dependencies remain queryable and searchable even after the source repository is deleted or moves.

Default location: ~/.drydock

Override with: - DRYDOCK_CATALOG environment variable - --catalog PATH flag on any catalog command

What Gets Recorded

When you ingest a repository:

  1. Repository metadata — name, source path, commit hash, branch, languages, file count
  2. Architectural facts — verdicts from principles (acyclic dependencies, stable abstractions, etc.) with evidence
  3. Components — discovered seams (contracts with multiple implementations) and clusters (cohesive modules)
  4. Capabilities — public symbols (classes, functions, constants) from every module, indexed for search
  5. Component code — optional storage; the extracted component tree is copied into the catalog so the library stays valid when source repos move

Re-ingesting the same repository at the same commit updates in place — no duplicates accumulate. A different commit is a separate entry. Once you forget a repository, all its records are removed (with forget requiring --yes).

Search and Discovery

Search is the point of the catalog. Use FTS5 to find symbols, component names, or facts:

# Find all symbols named "Parser"
drydock catalog search "Parser"

# Find repositories using TypeScript
drydock catalog search "TypeScript" --kind repo

# Find components with "plugin" in their contract
drydock catalog search "plugin" --kind component

# Find architectural facts about layering
drydock catalog search "layering" --kind fact

Identifiers are expanded into words. BaseParser is indexed as both BaseParser and Base Parser because FTS5 treats camelCase as a single token. A search for just parser would otherwise match nothing.

Commands

ingest — Record a new repository

drydock catalog ingest /path/to/project

Output:

{
  "repo_id": "repo_abc123def456",
  "components_recorded": 3,
  "components_stored": 0,
  "facts_recorded": 5,
  "capabilities_recorded": 42
}

Flags: - --store — Copy component code into the catalog (enables full code reconstruction, makes library portable) - --catalog PATH — Use a specific catalog root - --quiet, -q — Suppress progress output

list — Show all ingested repositories

drydock catalog list

Output (JSON):

[
  {
    "id": "repo_abc123def456",
    "name": "my-project",
    "source": "/path/to/project",
    "commit_sha": "abc123def456789",
    "branch": "main",
    "languages": ["python", "javascript"],
    "file_count": 142
  }
]

Output (markdown):

my-project                                        python javascript
  /path/to/project

Flags: - --markdown, -m — Human-readable format (default: JSON)

show — Details of one repository

drydock catalog show repo_abc123def456
# or by name:
drydock catalog show my-project

Returns repository metadata, all components found, and architectural facts with verdicts.

Flags: - --markdown, -m — Human-readable format

components — List extracted and detected components

drydock catalog components

Filters: - --repo ID — Only components from this repository - --kind seam|cluster|directory|manual — Filter by component type - --lang LANG — Only components with this language - --min-cohesion SCORE — Only components with cohesion ≥ SCORE (0.0–1.0) - --stored — Only components whose code is stored in the catalog

Output (JSON):

[
  {
    "id": "comp_abc123",
    "repo_id": "repo_xyz789",
    "name": "core",
    "kind": "seam",
    "contract": "IProcessor",
    "cohesion": 0.87,
    "languages": ["python"],
    "file_count": 5,
    "is_stored": false,
    "summary": "Core processing engine"
  }
]

search — Full-text search over the entire catalog

drydock catalog search "query"

Returns capabilities (symbols), components, repositories, and facts matching the query, ranked by relevance.

Filters: - --kind repo|component|capability|fact — Search only this subject type - --limit N — Maximum results to return (default: 25)

Output (JSON):

[
  {
    "subject_id": "cap_abc123",
    "subject_kind": "capability",
    "repo_id": "repo_xyz789",
    "title": "IProcessor",
    "snippet": "IProcessor Protocol core/processor.py interface",
    "rank": 3.14
  }
]

stats — Catalog statistics

drydock catalog stats

Output (JSON):

{
  "root": "/home/user/.drydock",
  "schema_version": 2,
  "repos": 5,
  "components": 18,
  "components_stored": 3,
  "facts": 25,
  "capabilities": 156,
  "languages": ["python", "typescript", "go"]
}

Output (markdown):

Catalog: /home/user/.drydock
Schema version: 2

Repositories: 5
Components: 18
  Stored: 3
Facts: 25
Capabilities: 156
Languages: python, typescript, go

diagram — Generate architectural diagrams

drydock catalog diagram repo_abc123 --kind layers

Generates a Mermaid diagram of the repository's dependency structure.

Diagram kinds: - layers (default) — Topological levels showing dependency flow - packages — Package-level coupling and stability metrics - seams — Seam contracts and their implementations

Output (JSON):

{
  "repo_id": "repo_abc123",
  "diagram_kind": "layers",
  "mermaid": "graph TD\n  layer0[...]\n  ..."
}

Output (markdown): Direct Mermaid source, viewable in GitHub Markdown or rendered via mermaid-cli:

graph TD
  layer0["Depth 0 (2 modules)"]
    core_models["core/models.py"]
  layer1["Depth 1 (3 modules)"]
    core_db["core/db.py"]
    ...

forget — Remove a repository from the catalog

drydock catalog forget repo_abc123 --yes

Prints what will be removed before deletion. Requires --yes to proceed.

What gets removed: - Repository metadata - All components from this repository - All facts about this repository - All capabilities (symbols) from modules in this repository - Optionally, stored component code (unless --keep-content)

Output (JSON):

{
  "repo_id": "repo_abc123",
  "repo_name": "my-project",
  "components_removed": 3,
  "content_trees_deleted": 1
}

Flags: - --yes — Skip confirmation - --keep-content — Preserve stored component code even if records are deleted - --markdown, -m — Confirm via stderr (for scripts)

Component Kinds

Kind When Example
seam Contract with 2+ implementations IProcessor with FastProcessor and SlowProcessor
cluster Modules with high internal coupling, external boundary Core business logic grouped together
directory All modules under one directory All files in src/utils/
manual Hand-annotated via plugin markers Explicitly marked with @plugin decorators

Python API

For programmatic access (TUI, GUI, agents), the Catalog class is the shared API:

from drydock.catalog.store import Catalog

catalog = Catalog()  # Uses ~/.drydock by default

# Record a repository
repo = catalog.put_repo(
    name="my-project",
    source="/path/to/project",
    source_kind="path",
    commit_sha="abc123",
    languages=["python"],
    file_count=42,
)

# Record a component
component = catalog.put_component(
    repo_id=repo.id,
    name="core",
    kind="seam",
    contract="IProcessor",
    file_count=5,
    languages=["python"],
    summary="Core processing module",
)

# Record architectural facts
fact = catalog.put_fact(
    repo_id=repo.id,
    principle="acyclic-dependencies",
    name="Acyclic Dependencies",
    verdict="holds",
    summary="No circular dependencies at load time.",
    evidence={"modules_total": 42, "cycles_found": 0},
)

# Record public symbols (capabilities)
cap = catalog.put_capabilities(
    component_id=component.id,
    symbols=[
        {"name": "IProcessor", "kind": "interface", "lineno": 5},
        {"name": "process", "kind": "method", "lineno": 12},
    ],
)

# Store component code in the catalog
stored_path = catalog.store_component_content(
    component.id,
    "/path/to/extracted/component",
)

# Search
hits = catalog.search("processor", kind="capability", limit=10)

# Query
repos = catalog.repos()
repo = catalog.repo("repo_abc123")

components = catalog.components(repo_id=repo.id, kind="seam")
component = catalog.component("comp_abc123")

facts = catalog.facts(repo_id=repo.id)
capabilities = catalog.capabilities(component_id=component.id)

# Statistics
stats = catalog.stats()

# Cleanup
catalog.forget(repo.id, delete_content=True)

catalog.close()

Queries return plain dataclasses, not database handles, so results can be held, serialized as JSON, or passed to other tools without database connections.

MCP Tools

The five catalog tools are available to AI assistants via MCP. Unlike analysis tools, drydock_catalog_ingest writes to the user's library; the others are read-only.

drydock_catalog_ingest

Ingest a project into the user's catalog.

{
  "name": "drydock_catalog_ingest",
  "description": "Ingest a repository into the user's component catalog",
  "inputSchema": {
    "type": "object",
    "properties": {
      "project_path": { "type": "string", "description": "Path to the project" },
      "store": { "type": "boolean", "description": "Store component code trees (default: false)" }
    },
    "required": ["project_path"]
  }
}

drydock_catalog_list

List all ingested repositories in the catalog.

{
  "name": "drydock_catalog_list",
  "description": "List all repositories in the catalog"
}

drydock_catalog_components

List components from the catalog, with optional filters.

{
  "name": "drydock_catalog_components",
  "description": "List components in the catalog",
  "inputSchema": {
    "type": "object",
    "properties": {
      "repo_id": { "type": "string", "description": "Filter by repository ID" },
      "kind": { "type": "string", "description": "Filter by kind (seam, cluster, directory, manual)" },
      "language": { "type": "string", "description": "Filter by language" },
      "min_cohesion": { "type": "number", "description": "Minimum cohesion score (0.0-1.0)" },
      "stored_only": { "type": "boolean", "description": "Only show stored components" }
    }
  }
}

drydock_catalog_facts

Retrieve architectural facts from ingested repositories.

{
  "name": "drydock_catalog_facts",
  "description": "Get architectural facts about repositories",
  "inputSchema": {
    "type": "object",
    "properties": {
      "repo_id": { "type": "string", "description": "Filter by repository ID" },
      "principle": { "type": "string", "description": "Filter by principle (e.g., 'acyclic-dependencies')" }
    }
  }
}

Search the catalog for symbols, components, repositories, and facts.

{
  "name": "drydock_catalog_search",
  "description": "Search the catalog",
  "inputSchema": {
    "type": "object",
    "properties": {
      "query": { "type": "string", "description": "Search query" },
      "kind": { "type": "string", "description": "Filter by subject kind (repo, component, capability, fact)" },
      "limit": { "type": "integer", "description": "Maximum results (default: 25)" }
    },
    "required": ["query"]
  }
}

Storage Schema

The catalog uses SQLite 3 with schema version 2. All queries use WAL mode for concurrent read-write access (a GUI can browse while the CLI is ingesting).

  • repos — Ingested projects with metadata
  • components — Detected and extracted seams and clusters
  • facts — Architectural principles and their verdicts
  • capabilities — Indexed symbols (classes, functions, constants) from modules
  • search — FTS5 index over the above

Content trees are stored as directories in ~/.drydock/components/<component_id>/.

Best Practices

  1. Ingest once, search forever — Run ingest once per project/commit. Querying doesn't require the source repository.

  2. Store components for portability — Use --store to copy extracted trees into the catalog. The library survives source repo deletion or relocation.

  3. Search to avoid duplication — Before building a new component, search: "Do I already have a JSON parser?" Identifiers expand into words so parseJSON matches a search for json.

  4. Use the Python API for automation — Queries return plain dataclasses, not database handles. Serializable, composable, and testable.

  5. Version your catalog — DRYDOCK_CATALOG makes it easy to maintain multiple libraries (one per team, one for archived projects, etc.).

See Also