feat: implement local code scout, training pipeline, and MCP tools

This commit is contained in:
emil28092005
2026-09-16 03:57:09 +03:00
parent 2e5ab98c56
commit ba24db2be5
33 changed files with 4347 additions and 27 deletions
+43
View File
@@ -0,0 +1,43 @@
# Dataset card: CodeSearchNet Python subset v1
## Selection rationale
The original **CodeSearchNet** project was produced by GitHub and Microsoft Research. It supplies function/documentation pairs, source repository identities, and code URLs. The original partitioning separates repositories. This gives a traceable starting point for a code retrieval experiment. [Original project](https://github.com/github/CodeSearchNet)
The local run uses the Parquet conversion hosted at [`code-search-net/code_search_net`](https://huggingface.co/datasets/code-search-net/code_search_net), pinned to revision `bd0cf261e357a3eb5c8fba490d23ec1a1cd59555`. The Hugging Face API reported **32,478 downloads and 337 likes** when inspected on September 16, 2026. These are popularity indicators, not accuracy or cleanliness guarantees.
The source dataset is established and attributable, but dates from an older Python ecosystem. Documentation comments are proxy labels; they do not represent the full distribution of coding-agent requests. We prefer this traceable source over an unexplained synthetic collection for the first baseline, then apply our own checks.
The [MiniLM base model](https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2) is published by Sentence Transformers under Apache 2.0. The same API inspection reported 254,208,155 downloads and 5,993 likes. It is a general English embedding model; popularity does not establish code-search quality. Its exact revision is `1110a243fdf4706b3f48f1d95db1a4f5529b4d41`.
## Preparation
1. Preserve the upstream train/validation/test assignments.
2. Require a repository, source path, and GitHub code URL.
3. Keep queries of 4–80 words and bounded, parseable Python functions.
4. Use the first documentation paragraph as the query.
5. Remove all Python docstrings and comments from the code input while retaining executable string literals.
6. Remove exact token duplicates, normalized structural duplicates, and exact query duplicates from the selected data.
7. Exclude cross-split repository, code, structural, and query matches.
8. Sample with deterministic hash priority and cap each repository at 200 selected examples per split.
Evaluation examples are selected first; overlapping development and training candidates are discarded. The structural hash normalizes identifiers and literals. It is a conservative clone heuristic and may discard distinct functions with similar structure. It cannot guarantee removal of every fork, translated query, or semantic near-duplicate.
The target sample sizes are 30,000 training pairs, 2,000 validation pairs, and 3,000 test pairs. Actual counts, rejection counts, source SHA-256 hashes, prepared-file hashes, and the overlap audit are written to `manifest.json`. A partial preparation run never publishes a completed dataset directory.
The prepared v1 dataset reached those sizes, covering 6,819 training repositories, 397 validation repositories, and 444 test repositories. All 15 pairwise overlap checks passed. The training source contained 412,178 rows; 45,464 failed the quality filters, and 846 additional candidates matched selected evaluation data by structural or query hash. A 12-example training-only spot check found plausible description/function pairs, including networking, file handling, rendering, and configuration. This small review is not a measured label-accuracy estimate.
## Provenance and use
Each prepared example retains its upstream dataset revision, repository, path, URL, content fingerprints, and split. Raw source data is not committed to this repository. CodeSearchNet's project code license does not override the licenses of the underlying repositories; the original project describes per-repository license records. [Source licensing information](https://github.com/github/CodeSearchNet#licenses)
No private local repositories or user conversations are used for this training run. No examples are sent to hosted teacher models.
## Evaluation limits
- The labels identify a paired function, not every valid answer to the query.
- The candidate pool is the selected split, not an entire live repository or a universal code corpus.
- Query language is primarily English. Russian retrieval is not established.
- The base model's pretraining overlap with evaluation examples cannot be excluded.
- Passing the overlap audit establishes the implemented checks only; it does not prove complete independence of repository families.
- Retrieval quality does not establish usefulness to Astra until a paired downstream evaluation is run.
+82
View File
@@ -0,0 +1,82 @@
# Training and evaluation
## What is trained
Version 0.1 fine-tunes `sentence-transformers/all-MiniLM-L6-v2` as a **shared bi-encoder**: the same small transformer maps queries and code into normalized vectors. This is supervised contrastive training, not training a language model from scratch and not reinforcement learning.
For each batch, the query's paired function is the labeled positive. Other functions in the batch are treated as negatives. The loss averages query-to-code and code-to-query cross-entropy. A temperature scales cosine scores. The model learns ranking without generating text.
At inference, the repository's vectors are already in the index. Each request encodes only its query, compares vectors, and combines rankings with BM25. This makes persistent inference practical without Ollama or vLLM.
## Reproduce the local run
From the repository root:
```bash
uv sync --extra train --extra mcp --extra dev --python 3.12
uv run --no-sync python -m micro_scout.download --output data/source
uv run --no-sync python -m micro_scout.data --source data/source --output data/csn-python-v1
uv run --no-sync python -m micro_scout.train \
--data data/csn-python-v1 --config configs/laptop.json \
--output runs/minilm-v1 --device cuda
```
The download helper pins both upstream revisions. Preparation publishes its output only after all three splits pass an overlap audit. It refuses to replace an existing dataset; use a new directory for a changed preprocessing experiment.
The laptop configuration uses 256 code tokens, 96 query tokens, a batch of 24, two epochs, AdamW at `2e-5`, gradient clipping, and mixed precision on CUDA. Full weights are trainable. A 100-minute training limit provides a checkpointed stopping point. Hardware-dependent duration must be measured.
Dataset shuffling is seeded. Exact bitwise reproducibility across PyTorch versions, GPUs, and kernels is not promised. `run.json`, the dataset manifest, and the saved configuration describe the actual experiment.
## Checkpoints and interruption
- `best/`: the checkpoint selected by validation MRR.
- `last/`: the latest checkpoint and optimizer, scheduler, scaler, and RNG state.
- `training.jsonl`: losses, validation scores, timing, and peak allocated GPU memory.
- `result.json`: final run metadata and completion state.
Training checks a stop flag between batches after `SIGINT` or `SIGTERM`, then validates and saves. For a normal stopped run:
```bash
uv run --no-sync python -m micro_scout.train \
--data data/csn-python-v1 --config configs/laptop.json \
--output runs/minilm-v1 --resume runs/minilm-v1/last --device cuda
```
Resume requires the same training configuration and dataset manifest. This command is intended for the project's own local optimizer states. Model weights use safetensors, and remote custom model code is disabled.
## Evaluation protocol
Checkpoint selection uses 512 fixed validation pairs. The complete validation set is evaluated separately. The final test set is reserved until the model and retrieval settings are frozen.
Every query ranks the same full candidate pool from its split. The comparison includes:
1. BM25 with identifier-aware tokenization.
2. The original pretrained MiniLM encoder.
3. The locally fine-tuned encoder.
4. Hybrid retrieval for each encoder, using the same fixed reciprocal-rank fusion.
All methods receive code without its documentation query. Report MRR, MRR@10, recall@1/5/10, and the candidate count. A paired bootstrap resamples whole repositories to estimate uncertainty in MRR differences while retaining within-repository correlation. Repository-macro MRR is also reported so large projects do not hide performance on smaller ones.
```bash
uv run --no-sync python -m micro_scout.evaluate --split validation --device cuda
uv run --no-sync python -m micro_scout.evaluate --split test --device cuda
```
The metrics and per-query ranks are saved under `runs/minilm-v1/evaluation/`. The evaluator never updates weights. The final test must not become a repeated hyperparameter-selection loop.
## Latency
```bash
OPENBLAS_NUM_THREADS=1 uv run --no-sync micro-scout benchmark \
--index .micro-scout/index.sqlite --model runs/minilm-v1/best \
--device cpu --threads 4 --iterations 100 \
--output runs/minilm-v1/latency.json
```
This measures the warm harness, including search, context assembly, and checking returned files. Startup is reported separately. Five repeated development queries are used; these are not representative production traffic. Index construction is also reported separately.
## What these results cannot establish
CodeSearchNet descriptions are weak task labels. A function may have several valid alternatives, while the metric assumes only one positive. The test does not measure multi-file reasoning, patch correctness, context completeness, or Astra's success rate. A follow-up experiment should compare a fixed solver with and without micro-scout on real, held-out repository tasks.
Luna-based data generation and solver-feedback training are deferred. The first run requires no teacher API key or paid model calls.
+51
View File
@@ -0,0 +1,51 @@
# CLI and MCP usage
## Tool behavior
`scout_search` accepts a natural-language query, `top_k` from 1 to 50, and `max_chars` from 100 to 100,000. It returns ranked snippets and up to three graph neighbors within the shared source-character budget. Overlapping source lines are not repeated. Scores are not probabilities.
Markdown is excluded from search by default; set `include_docs=true` to include it. An optional `language` filter narrows candidates before rank fusion, for example to `python` or `typescript`. CLI equivalents are `--include-docs` and `--language python`.
Each snippet includes an opaque symbol ID, repository-relative path, line range, source text, file SHA-256, and verification/truncation flags. The response also identifies the index snapshot and model fingerprint. Truncation occurs at line boundaries.
`scout_read` takes an ID returned by the current index. It checks the source hash again before returning a larger range. It accepts no arbitrary filesystem path.
`scout_status` describes the loaded index. `scout_feedback` is available only with an explicit trace path; it accepts IDs actually returned by one of the last 1,000 searches in the current process. Feedback is stored locally and never updates serving weights.
## Example MCP host configuration
Use the executable in the environment where the package was installed. Replace all paths:
```json
{
"mcpServers": {
"micro-scout": {
"command": "/absolute/path/to/micro-scout/.venv/bin/micro-scout",
"args": [
"serve",
"--index", "/absolute/path/to/project/.micro-scout/index.sqlite",
"--model", "/absolute/path/to/micro-scout/runs/minilm-v1/best",
"--device", "cpu",
"--trace", "/absolute/path/to/micro-scout/runs/session.jsonl"
],
"env": {"OPENBLAS_NUM_THREADS": "1"}
}
}
}
```
The server uses the official [MCP Python SDK v1 maintenance line](https://py.sdk.modelcontextprotocol.io/v1/), pinned below v2. It speaks stdio and does not open a network listener. The weights and index stay resident for the server's lifetime. A compatible host can call these tools; this alone does not establish downstream model quality.
## Files and index updates
Indexing honors Git's ignored-file rules when available and skips hidden directories, common dependency/build directories, symlinks, unsupported extensions, and files above 1 MB. Python symbols use AST boundaries. Other supported text/code formats use 60-line chunks. Incomplete Python falls back to chunks.
An index is one SQLite file, atomically replaced after a successful build. A model fingerprint prevents combining incompatible query and code embeddings. Reindexing with unchanged weights reuses embeddings for unchanged normalized code.
Returned files are checked against their indexed content hashes. Modified or missing files are excluded with warnings; reads of stale IDs fail. New files require reindexing. This is a snapshot workflow, not a filesystem watcher. Restart the server after rebuilding an index so it loads the new snapshot.
## Trace contents
With `--trace`, the local JSONL trace records queries, returned IDs, snapshot, latency, and feedback. Source snippets are not copied into search trace records. Without this option, query traces are not written. The repository's default `runs/` directory is excluded from Git.
The tool never executes indexed code. Source text is untrusted input for the consuming model and should be treated as data, including any instruction-like comments it contains.