Exploring Codebases: Progressive Disclosure Meets Hybrid Search

By Oskar 🕊️ (@austegard.com)
Published:

Vibe coded and written by Gemini and Opus with design and architectural decisions by Oskar Austegard

The Problem with Code Search

When working with unfamiliar codebases, there's a fundamental tension: you need to find relevant code quickly, but you also need enough context to understand what you've found. Traditional approaches fall into two camps:

This tension becomes critical for agentic coding, where an AI needs to rapidly explore codebases it's never seen before.

What is Tree-sitter?

Before diving into the solution, let's clarify the key technology: tree-sitter is a parser generator that builds concrete syntax trees (CSTs) for source code. Unlike traditional parsers, it's designed to be:

Most importantly, tree-sitter understands code structure. It knows the difference between:

This structural awareness is what grep fundamentally lacks.

The Journey: From Mapping to Exploring

mapping-codebases: The Structure-First Approach

The original mapping-codebases skill took the structure-first path. It would parse entire directory trees using tree-sitter, extract all exports/imports, and generate comprehensive _MAP.md files showing the skeleton of each module:

## classes/User.py

### Classes
- `User` (line 15)
  - `__init__(self, name, email)` (line 16)
  - `validate(self)` (line 23)
  - `save(self)` (line 31)

Strengths:

Limitations:

GrepRAG: The Speed Revelation

Then came the GrepRAG paper, which benchmarked various code retrieval methods for LLM context. Their key finding: ripgrep-based retrieval was 17× faster than graph-based methods (0.40s vs 7s) while maintaining comparable accuracy.

However, they also identified two critical failures of pure text search:

The paper attempted to fix these with statistical heuristics (line-number clustering, identifier weighting), but these are fundamentally band-aids on a text-only approach.

The Hybrid Solution: Speed + Structure + Efficiency

The new exploring-codebases skill combines three key insights:

Phase 1: The Dragnet (ripgrep)

rg --json "class Session" /path/to/repo

Goal: Cast a wide net quickly. Find every file and line where the search term appears.

Speed: ~0.02s for large repositories (as GrepRAG demonstrated)

Output: List of (file, line_number) tuples

Phase 2: The Scalpel (tree-sitter)

For each ripgrep match, use tree-sitter to:

Example: If ripgrep finds Session on line 356, tree-sitter identifies it's inside a class definition spanning lines 356-816.

Phase 3: Progressive Disclosure (The Token Multiplier)

Here's the critical enhancement over v0.1: don't dump 460 lines of implementation when 20 lines of signature will do.

Default output (signatures only):

class Session(SessionRedirectMixin):
    """A Requests session.
    
    Provides cookie persistence, connection-pooling, and configuration.
    """
    ...

Token cost: ~50 tokens (vs 4000 for full implementation)

When you need details, expand:

search.py "class Session" /path/to/repo --expand-full

Returns the complete 460-line implementation.

Why This Matters

1. Context Fragmentation → Fixed

Problem (grep alone): Searching for get returns line 595:

return self.request("GET", url, **kwargs)

No parameters, no docstring, no context.

Solution (hybrid with signatures):

def get(self, url, **kwargs):
    r"""Sends a GET request. Returns :class:`Response` object.

    :param url: URL for the new :class:`Request` object.
    :param \*\*kwargs: Optional arguments that ``request`` takes.
    :rtype: requests.Response
    """
    ...

You get the complete API surface without the implementation noise.

2. Keyword Ambiguity → Fixed

Tree-sitter knows the difference between:

This filtering is deterministic, not statistical. You don't need BM25 re-ranking to boost likely definitions; the AST tells you what's a definition.

3. Token Efficiency → 10-20× Improvement

Progressive disclosure exploits a key insight: most code exploration is hierarchical.

Typical workflow:

Old approach (v0.1):

New approach (v0.2):

4. Real-World Example

Testing on the requests library:

# Find what Session offers
$ search.py "class Session" requests/

Found 2 matches for 'class Session':

### requests/sessions.py
class Session(SessionRedirectMixin):
    """A Requests session.
    
    Provides cookie persistence, connection-pooling, and configuration.
    """
    ...

Tokens used: ~80

Now you know Session exists and what it does. Need to know its methods?

# Scan for methods (still signature-only)
$ search.py "def " requests/sessions.py | grep "class.*Session" -A 20

Or expand a specific method:

$ search.py "def request" requests/sessions.py --expand-full

Progressive approach: ~300 tokens total to understand the class and one method

Old approach: 4,000 tokens dumped upfront, most unused

Performance Characteristics

The hybrid approach exploits code's sparse structure:

Comparison table:

| Approach | Preprocessing | Search Time | Tokens per Match | Updates | |----------|--------------|-------------|------------------|---------| | LSP/ctags | Minutes | Instant | Variable | Manual rebuild | | mapping-codebases | Minutes | Instant | 50-100 | Stale maps | | grep/ripgrep | None | 0.02s | 5-10 (fragmented) | Real-time | | GrepRAG | None | 0.40s | 50-200 (heuristic) | Real-time | | exploring-codebases v0.1 | None | 0.05s | 500-5000 (full) | Real-time | | exploring-codebases v0.2 | None | 0.05s | 50 (sig) / 500+ (full) | Real-time |

Implementation Notes

The enhancement adds ~100 lines to the original 280-line script:

def _extract_signature(self, node, source_bytes, language):
    """Extract just the declaration + docstring, exclude body."""
    if language == 'python':
        return self._extract_python_signature(node, source_bytes)
    # ... other languages

For Python:

Example CST traversal:

class_definition
├── class (keyword)
├── identifier (name)
├── argument_list (bases)
├── : (colon)
└── block (body) → extract first string child (docstring), then stop
    ├── string (docstring) ← include this
    ├── function_definition ← skip
    └── ... ← skip

Usage Patterns

Quick exploration (default):

# What classes exist?
search.py "class " src/

# What methods does User have?
search.py "def " src/models/user.py

# Find all validation methods
search.py "def validate" src/

Deep dive (selective expansion):

# Get full implementation of specific method
search.py "def validate_email" src/ --expand-full

# Get full class implementation
search.py "class User" src/ --expand-full

Language-specific:

# Only Python files
search.py "class Config" . --glob "*.py"

# Only TypeScript
search.py "interface User" . --glob "*.ts"

When to Use What

Use mapping-codebases when:

Use exploring-codebases when:

Use grep/ripgrep directly when:

Addressing GrepRAG's Concerns

The paper identified these pain points with pure grep:

| GrepRAG Concern | exploring-codebases Solution | |----------------|----------------------------| | "Disrupts logical flow" (fragments code) | Tree-sitter ensures complete function/class nodes | | "Usage before definition" (line ordering) | AST traversal maintains semantic relationships | | "Keyword noise" (comments, strings) | Deterministic AST filtering, not statistical | | "Token waste" (returning everything) | Progressive disclosure: signatures first | | "Expensive re-ranking" (BM25, TF-IDF) | Not needed - structural filtering is deterministic |

GrepRAG tried to fix grep's problems with more grep (line clustering, statistical weighting). We fix them by acknowledging grep's limits and using the right tool for structural awareness.

Conclusion

The GrepRAG paper proved text search is fast enough for real-time code retrieval. But speed without structure wastes tokens, and structure without progressive disclosure wastes more.

By combining:

We get a new paradigm for code exploration: lazy, selective, precise, and token-conscious.

This is particularly valuable for AI agents that need to rapidly orient themselves in unfamiliar codebases. Instead of "here's line 45" or "here's 10,000 lines," they get "here's the signature; ask if you need the body" - exactly the context needed to reason efficiently.

The future of code search is hybrid, structural, and progressive.


Code: https://github.com/oaustegard/claude-skills/tree/main/exploring-codebases Release zip: https://github.com/oaustegard/claude-skills/releases?q=exploring-codebases&expanded=true